Multi-Region Disaster Recovery Test Checklist
Retrospective: this article looks back at events from December 2021, written in 2026 with the benefit of hindsight.
Use this checklist to plan and run a multi-region disaster recovery test for an AWS workload.
Before the test
- Workload owner, RTO and RPO targets documented.
- DR strategy chosen (backup/restore, pilot light, warm standby, active-active).
- Data replication status verified (replication lag, last successful backup copy).
- Infrastructure as code for the secondary region is current.
- KMS keys, secrets and parameters exist in the secondary region.
- DNS failover mechanism documented and tested in a non-production domain.
- Test scope, success criteria and rollback plan agreed.
- Stakeholders and customers informed if user-facing impact is possible.
During the test
- Start time recorded.
- Failover triggered using the documented procedure, by someone other than its author.
- Application health checks pass in the secondary region.
- Data integrity verified (recent transactions present).
- Time to restore service recorded.
- Monitoring and alerting work in the secondary region.
- Support and communication steps exercised.
Failback
- Failback procedure executed and timed.
- Data changes made during failover reconciled.
After the test
- Actual RTO and RPO compared with targets.
- Gaps documented with owners and dates.
- Runbook updated.
- Results reported to leadership.
Frequency
- Critical workloads: at least annually; after major architecture changes.
- The AWS us-east-1 Outage of December 2021: When the Control Plane Fails Incident Teardowns
- How to Plan Multi-Region Failover for Critical AWS Workloads How-To & Hardening
- CIO Brief: Concentration Risk in a Single Cloud Region CIO Briefings