The AWS S3 Outage of February 2017: A Typo That Broke the Internet
Retrospective: this article looks back at events from February 2017, written in 2026 with the benefit of hindsight.
On February 28, 2017, Amazon S3 in the US-EAST-1 region became unavailable for about four hours. Thousands of websites and apps stopped working, and even the AWS Service Health Dashboard struggled to report the outage because it depended on S3 too.
What happened
According to AWS's post-incident summary, an engineer debugging a billing system ran a command intended to remove a small number of servers. An input was entered incorrectly and a much larger set of servers was removed, including servers supporting two critical S3 subsystems. Restarting those subsystems at that scale took far longer than expected.
Why it mattered
S3 had become a foundation for much of the internet: static websites, images, application assets, backups and data pipelines. Many architectures assumed it would simply always be there. Applications built entirely in one region, with no fallback, went down with it.
Lessons in hindsight
- Regions fail. Rarely, but they do. Critical workloads need a plan for a regional outage, even if that plan is a degraded mode rather than full failover.
- Know your hidden dependencies. The status dashboard depending on the service it monitors was an example everyone remembered.
- Guard rails on powerful commands. AWS changed its tooling to remove capacity more slowly and to block removals below a minimum safe level. Your own automation deserves the same protections.
- Availability is a security property. Confidentiality, integrity and availability are the classic triad for a reason.
Ten years on
Cross-region replication, multi-region access points and better guidance on resilience make regional fallback far easier than in 2017. Yet major regional outages continued, including in US-EAST-1 in 2021 and 2025. The question for every team remains: what happens to our business if one region disappears for a day?
- How to Architect S3 Workloads for Multi-Region Resilience How-To & Hardening
- Running a Cloud Outage Tabletop Exercise for Your IT Team How-To & Hardening
- CIO Brief: Availability Is Security — Building Outage Risk Into Your Cloud Strategy CIO Briefings