Configuration drift and why it breaks cloud recovery
Configuration drift is the difference between the running environment and its intended definition. It breaks recovery in two places: production drifts away from the code or snapshot you would rebuild from, and the recovery environment drifts away from production. Drift detection and a change history are how you see both before an incident.
Ranking current as of September 2026 · By the Cloud Resilience Vendors research desk
What is configuration drift?
Drift is any difference between what is running in your cloud accounts and what you believe is running, as written in infrastructure as code or captured in the last configuration snapshot. It usually comes from changes made by hand in a console, from scripts and automation that act outside the normal pipeline, and from emergency fixes that nobody writes back into code.
Drift is normal in a live estate. It becomes a recovery problem only when nobody sees it, because every rebuild starts from a definition, and a definition that no longer matches production rebuilds something slightly different.
Where does drift hurt a recovery?
- Production drifts from its definition. If a security group rule, an IAM policy or a DNS record was changed by hand and never captured, a rebuild from code or from an older snapshot brings back the old version. The application may start and still fail, because one permission or route is missing.
- The recovery environment drifts from production. In pilot light and warm standby strategies, a copy of the infrastructure already exists in another region. AWS's guidance on testing cloud disaster recovery tells teams to manage configuration drift in the recovery region, so that its infrastructure, data and configuration match what production needs, and to check machine images and service quotas there.
Both kinds are silent until the day you need the rebuild, which is why they are best found by a tool that compares continuously rather than by a yearly test.
What does good drift detection look like?
- It compares against a known-good state, not only against the last change.
- It covers resources that were never in code, not only the ones that were.
- It records what changed, when and, where the platform exposes it, who made the change, so you can pick a point in time to restore to.
- It alerts in a way that reaches the owner of the service, not only a central team.
Which tools in the ranking publish drift features?
Firefly documents drift detection, mutation tracking through its Event Center and a configuration history. ControlMonkey lists drift detection with granular alerting and a scanner for changes made in the console. HCP Terraform runs drift detection through health assessments, in its Standard and Premium editions only, and only for resources already in Terraform. StackGuardian lists continuous drift detection with audit reports. Arpio and Cohesity do not publish drift detection or a change history on the pages we reviewed, and Veeam does not offer it. The drift and change history scores for each vendor are in the ranking and on each vendor profile.
What should you take from this lesson?
Ask two questions of any recovery plan: when did we last compare production with the definition we would rebuild from, and when did we last compare the recovery environment with production? If the answer is "during the last test", drift has had months to build up.
Sources
- AWS whitepaper, testing recovery: https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/testing-disaster-recovery.html
- Firefly docs, Backup and DR overview: https://docs.firefly.ai/detailed-guides/backup-and-disaster-recovery.md
- ControlMonkey, Home: https://controlmonkey.io/
- HashiCorp docs, Health assessments: https://developer.hashicorp.com/terraform/cloud-docs/workspaces/health
- StackGuardian, Home: https://www.stackguardian.io/