Multi-CloudIncident TeardownsRetrospectives

The CrowdStrike Outage (July 19, 2024): 8.5 Million Windows Machines Down

By OnCloudSec Research Team · Published Oct 6, 2026 · 5 min read

Facts in this article were checked against the sources listed below as of Oct 5, 2026.

Retrospective: this article looks back at events from July 2024, written in 2026 with the benefit of hindsight.

On July 19, 2024, Windows computers around the world began crashing with blue screens and failing to restart. Airlines grounded flights, hospitals postponed procedures, and banks and broadcasters went offline. Microsoft estimated that about 8.5 million Windows devices were affected.

How it unfolded
  1. Content update releasedJuly 19, 2024 at 04:09 UTC: CrowdStrike releases an update to Channel File 291 for its Falcon Windows sensor.
  2. Input mismatchThe update expects 21 input fields; the sensor code supplies 20, triggering an out-of-bounds memory read.
  3. Kernel crash loopWindows hosts crash and fail to boot because the sensor loads early in startup.
  4. Reverted in 78 minutesCrowdStrike reverts the update at 05:27 UTC — but affected machines need manual repair.
  5. Slow recoveryBitLocker recovery keys and hands-on fixes delay restoration across airlines, hospitals and banks.

The cause wasn't a cyberattack. It was a routine content update to CrowdStrike Falcon, one of the most widely used endpoint security products. CISA quickly confirmed the outage was not a malicious incident — and warned that attackers were already exploiting the confusion with phishing.

What happened

CrowdStrike released an update to a configuration file known as Channel File 291 at 04:09 UTC and reverted it at 05:27 UTC — a distribution window of 78 minutes. According to CrowdStrike's root cause analysis published in August, the update used a template that defined 21 input fields, while the sensor code supplied only 20. The mismatch caused an out-of-bounds memory read in the sensor, which runs in the Windows kernel. Windows crashed.

Because the sensor loads early in the boot process, affected machines crashed again on every restart. Reverting the file centrally didn't help machines that couldn't stay up long enough to receive the fix. Each needed manual repair: booting into Safe Mode or the Windows Recovery Environment and deleting the faulty file.

Why recovery took so long

  • BitLocker. Many affected devices were encrypted. Before repairing them, technicians needed each device's BitLocker recovery key — and many organizations struggled to retrieve thousands of keys quickly, sometimes because the systems that held them were also down.
  • Hands-on work. Remote workers' laptops, kiosks and point-of-sale systems needed someone physically present.
  • Cloud VMs. Windows virtual machines in Azure and AWS running Falcon were affected too. Microsoft and AWS published recovery guidance, including attaching OS disks to repair VMs, restoring from snapshots and using serial consoles.

Why it mattered

Security tools are a resilience risk. Software that runs with kernel-level privileges on every device can take every device down at once. The same applies to other universal software: operating system updates, management agents and drivers.

Availability is part of security. The impact looked like a ransomware attack — company-wide outage, manual recovery — without an attacker.

Recovery readiness matters as much as prevention. Organizations that could retrieve recovery keys at scale and had tested mass-recovery procedures restored service faster.

What changed afterward

CrowdStrike introduced staged rollouts and more customer control over content updates. At Ignite 2024, Microsoft announced the Windows Resiliency Initiative, including plans to let security vendors run outside the kernel and Quick Machine Recovery to remotely fix devices that can't boot.

What to do now

  1. Make BitLocker recovery keys retrievable at scale. Confirm keys for Entra-joined and hybrid-joined devices are escrowed to Entra ID and visible in Intune, and that AD-joined devices back up to Active Directory. Plan how help desk staff can retrieve keys if normal tools are unavailable, and consider self-service key retrieval for users.
  2. Ask vendors about staged updates. For every agent that runs on all devices, find out whether sensor versions and content updates can be staged in rings, and how fast the vendor can roll back.
  3. Define update rings. IT test devices first, then early adopters, general population, and critical systems such as servers and point-of-sale last. Where vendors allow, keep critical systems one version behind.
  4. Document mass-recovery runbooks. Bootable recovery media, scripts for common fixes, priority order for critical roles, and procedures for Azure and AWS VMs (snapshots, repair VMs, serial console access).
  5. Keep an out-of-band communication channel. If company laptops are down, how will you reach employees?
  6. Back up before major updates for critical servers, and keep recent snapshots for cloud VMs.
  7. Include vendor-caused outages in contracts. Review quality assurance, notification and liability terms with critical software vendors.

How to detect a mass failure early

  • Device health signals. Alert on spikes in devices going offline or failing health checks in Intune, Defender for Endpoint or your RMM tool.
  • Crash telemetry. Sudden increases in system crashes reported through Windows error reporting or monitoring tools.
  • Vendor status pages. Subscribe to status notifications for every critical agent vendor.
  • Help desk volume. A sudden surge in "my computer won't start" tickets is itself a detection signal — route it to incident management quickly.

Common mistakes

  • Storing recovery keys only on systems that might go down.
  • Updating every device at once because the vendor makes it the default.
  • No plan for remote workers, whose devices may need shipping or guided recovery.
  • Treating it as a vendor problem only. Your recovery speed was yours to own.

Lessons for cloud workloads

Windows virtual machines in Azure and AWS were affected just like laptops, but recovery options differed. Teams that kept recent disk snapshots could restore quickly; teams that relied on golden images could redeploy stateless servers from scratch. Others attached the affected OS disk to a working repair VM to remove the faulty file — the az vm repair extension in Azure CLI automates parts of that process. Serial console access helped where it was already enabled.

The broader lesson for cloud teams is to design for the possibility that every instance of a given image fails at once: keep infrastructure as code current, prefer stateless and replaceable servers, take regular snapshots of stateful systems, and know which encryption keys and access paths you'll need during recovery.

Questions for leadership

  • If every Windows device and server failed to start tomorrow, how long would recovery take?
  • Can we retrieve thousands of encryption recovery keys within an hour?
  • Which of our vendors push updates to every device simultaneously, and can we stage them?

Key takeaways

  • A faulty CrowdStrike content update crashed about 8.5 million Windows devices on July 19, 2024.
  • The update was reverted in 78 minutes, but affected devices needed manual repair.
  • BitLocker key access and mass-recovery readiness determined how fast organizations recovered.
  • Stage updates where possible and rehearse recovery from vendor-caused outages.

Sources

  1. CISA: Widespread IT outage due to CrowdStrike update
  2. TechTarget: CrowdStrike outage explained — what caused it and what's next
crowdstrike outage2024

More on this story