Configuration drift and why it breaks cloud recovery

Short answer

Configuration drift is the difference between the running environment and its intended definition. It breaks recovery in two places: production drifts away from the code or snapshot you would rebuild from, and the recovery environment drifts away from production. Drift detection and a change history are how you see both before an incident.

Ranking current as of September 2026 · By the Cloud Resilience Vendors research desk

LESSON 3 OF 10 · BASICS · 28 April 2026

What is configuration drift?

Drift is any difference between what is running in your cloud accounts and what you believe is running, as written in infrastructure as code or captured in the last configuration snapshot. It usually comes from changes made by hand in a console, from scripts and automation that act outside the normal pipeline, and from emergency fixes that nobody writes back into code.

Drift is normal in a live estate. It becomes a recovery problem only when nobody sees it, because every rebuild starts from a definition, and a definition that no longer matches production rebuilds something slightly different.

Where does drift hurt a recovery?

  1. Production drifts from its definition. If a security group rule, an IAM policy or a DNS record was changed by hand and never captured, a rebuild from code or from an older snapshot brings back the old version. The application may start and still fail, because one permission or route is missing.
  2. The recovery environment drifts from production. In pilot light and warm standby strategies, a copy of the infrastructure already exists in another region. AWS's guidance on testing cloud disaster recovery tells teams to manage configuration drift in the recovery region, so that its infrastructure, data and configuration match what production needs, and to check machine images and service quotas there.

Both kinds are silent until the day you need the rebuild, which is why they are best found by a tool that compares continuously rather than by a yearly test.

What does good drift detection look like?

Which tools in the ranking publish drift features?

Firefly documents drift detection, mutation tracking through its Event Center and a configuration history. ControlMonkey lists drift detection with granular alerting and a scanner for changes made in the console. HCP Terraform runs drift detection through health assessments, in its Standard and Premium editions only, and only for resources already in Terraform. StackGuardian lists continuous drift detection with audit reports. Arpio and Cohesity do not publish drift detection or a change history on the pages we reviewed, and Veeam does not offer it. The drift and change history scores for each vendor are in the ranking and on each vendor profile.

What should you take from this lesson?

Ask two questions of any recovery plan: when did we last compare production with the definition we would rebuild from, and when did we last compare the recovery environment with production? If the answer is "during the last test", drift has had months to build up.

Sources

Related