CIO Brief: Availability Is Security — Building Outage Risk Into Your Cloud Strategy
Retrospective: this article looks back at events from February 2017, written in 2026 with the benefit of hindsight.
The short version: In 2017, a mistyped command at Amazon took down a core storage service in one region for about four hours, and thousands of websites went with it. It was not a hack. It was still a business continuity event for every customer affected.
Why availability belongs in your security strategy
Security is usually discussed as keeping data private and accurate. Keeping systems available is the third part of the job, and outages — whether caused by attackers, bugs or human error — hit revenue and reputation the same way.
The business impact
- Lost revenue for every hour customers cannot transact.
- Contract penalties if you promise uptime to your own customers.
- Operational chaos when internal tools fail at the same time.
Questions to ask your team
- Which of our critical services run in only one cloud region?
- What would happen to customers if that region went down for a day?
- Have we decided, for each service, whether to fail over, degrade or wait?
- When did we last test it?
What good looks like
Critical services have a documented recovery approach matched to their business value. Not everything needs multi-region failover — that can double costs — but the most important services have a tested plan, and leadership has explicitly accepted the risk for the rest.
The decision
Ask for a short list of services ranked by the cost of a day of downtime, with the current recovery plan and an estimated cost to improve each. Then fund the top few. Outages like this one have repeated since 2017 — the companies that planned for them recovered fastest.
- The AWS S3 Outage of February 2017: A Typo That Broke the Internet Incident Teardowns
- How to Architect S3 Workloads for Multi-Region Resilience How-To & Hardening
- Running a Cloud Outage Tabletop Exercise for Your IT Team How-To & Hardening