Endpoint Agent Update Risk Checklist
Retrospective: this article looks back at events from July 2024, written in 2026 with the benefit of hindsight.
Security agents and other kernel-level software can take down entire fleets. Use this checklist to manage the risk of endpoint agent updates.
Vendor capabilities
- Does the vendor offer staged or ring-based rollout for sensor (agent) versions?
- Does it offer staged rollout or customer control for content/configuration updates (the type that caused the CrowdStrike outage)?
- What testing does the vendor perform before release?
- What is the vendor's rollback capability and speed?
Your deployment
- Update rings defined (for example: IT test group → early adopters → general → critical systems last).
- Critical systems (servers, point-of-sale, clinical devices) in later rings.
- Agent version policies set to N-1 for critical systems where appropriate.
Recovery readiness
- BitLocker recovery keys escrowed and retrievable at scale.
- Recovery media and scripts prepared.
- Remote recovery options for remote workers documented.
- Out-of-band communication channel for employees.
- Azure and AWS VM recovery procedures documented and tested.
Monitoring
- Alerts for spikes in device crashes or offline devices.
- Vendor status page and notifications subscribed.
Contracts
- Vendor obligations for quality assurance and incident communication reviewed.
- Liability and service credits understood.
Practice
- Mass endpoint failure scenario included in tabletop exercises.
- The CrowdStrike Outage (July 19, 2024): 8.5 Million Windows Machines Down Incident Teardowns
- How to Recover BitLocker-Protected Azure VMs and Endpoints at Scale How-To & Hardening
- CIO Brief: When Your Security Tool Causes the Outage CIO Briefings