AWSHow-To & HardeningRetrospectives

Running a Cloud Outage Tabletop Exercise for Your IT Team

By OnCloudSec Research Team · Published Oct 6, 2026 · 1 min read

Retrospective: this article looks back at events from February 2017, written in 2026 with the benefit of hindsight.

Outages are rare enough that teams forget how to handle them. A tabletop exercise — a structured discussion of a realistic scenario — is the cheapest way to find gaps before a real failure does.

Pick a realistic scenario

Start with something that has actually happened: "Our primary cloud region's object storage is unavailable for six hours, starting at 10 a.m. on a weekday." Variations: identity provider outage, DNS outage, a security tool update that crashes endpoints.

Who should attend

The on-call engineer, an application owner, someone from customer support, a communications lead and a business decision-maker. Ninety minutes is enough.

Run the exercise

  1. Detection (15 minutes): How would we find out? Monitoring, customers, social media?
  2. Triage (20 minutes): What is affected? Which services depend on the failed component — including hidden dependencies like status pages, ticketing or SSO?
  3. Response (25 minutes): What can we do? Fail over, degrade gracefully, wait? Who decides?
  4. Communication (15 minutes): What do we tell customers, staff and leadership, and through which channel if our usual tools are down?
  5. Recovery (15 minutes): How do we confirm everything is back and data is consistent?

Capture the gaps

Write down every "we don't know" and "that would depend on one person." Assign each gap an owner and a date.

Common findings

  • The status page or runbook is hosted on the same platform that failed.
  • Only one person knows how to trigger failover.
  • No pre-approved customer message templates.
  • Backups exist but restores have never been timed.

Repeat

Run a tabletop at least twice a year, rotating scenarios. Each one gets faster and more useful.

cloud outage tabletop exercise checklistAWS S3 us-east-1 outage2017

More on this story