How to Architect S3 Workloads for Multi-Region Resilience
Retrospective: this article looks back at events from February 2017, written in 2026 with the benefit of hindsight.
Amazon S3 is extremely durable, but a single region can still become unavailable. Here is how to design S3-backed workloads to keep running — or at least degrade gracefully — during a regional outage.
Step 1: Classify your buckets
Not every bucket needs multi-region resilience. Classify each one:
- Critical serving data: assets your application needs to respond to customers.
- Important data: backups and data you must not lose but can wait to read.
- Everything else: logs, scratch data, test buckets.
Step 2: Replicate critical data
Use S3 Cross-Region Replication (CRR) to copy objects to a bucket in a second region.
- Enable versioning on source and destination buckets (required).
- Create a replication rule, scoped to the prefixes that matter.
- Use S3 Replication Time Control if you need predictable replication times.
- Replicate encryption settings and confirm the destination KMS key policy allows it.
Step 3: Make reads failover-ready
- Serve static content through CloudFront with an origin group that fails over to the second region's bucket.
- For applications, use S3 Multi-Region Access Points to route requests to an available bucket.
- Avoid hardcoding region-specific bucket names in application code.
Step 4: Plan for writes
Writes during an outage are harder. Decide whether the application should queue writes, write to the secondary region, or pause non-essential features.
Step 5: Test it
Simulate a regional failure in a non-production environment by blocking access to the primary bucket and confirm the application keeps serving.
Common mistakes
- Replicating data but leaving IAM roles, KMS keys and DNS tied to one region.
- Forgetting replication only applies to new objects unless you run S3 Batch Replication for existing ones.
- Never testing failover.
- The AWS S3 Outage of February 2017: A Typo That Broke the Internet Incident Teardowns
- Running a Cloud Outage Tabletop Exercise for Your IT Team How-To & Hardening
- CIO Brief: Availability Is Security — Building Outage Risk Into Your Cloud Strategy CIO Briefings