Cloud disaster recovery strategies: from backup and restore to active/active

Short answer

AWS describes four cloud disaster recovery strategies: backup and restore, pilot light, warm standby and multi-site active/active. Each step up lowers recovery time and generally raises running cost. In every one, the infrastructure in the recovery location has to exist or be rebuilt, which is where configuration capture matters.

Ranking current as of September 2026 · By the Cloud Resilience Vendors research desk

LESSON 2 OF 10 · BASICS · 10 March 2026

What are the four strategies?

AWS's whitepaper on recovery options in the cloud describes them in order of increasing readiness:

  1. Backup and restore. Data is backed up and restored after an incident. AWS calls it suitable "for mitigating against data loss or corruption". Infrastructure, configuration and application code have to be redeployed, which gives it the longest recovery time.
  2. Pilot light. Data is replicated to another region and a copy of the core infrastructure is provisioned, with application servers "switched off" until testing or failover.
  3. Warm standby. A "scaled down, but fully functional, copy of your production environment in another Region" runs all the time.
  4. Multi-site active/active. The workload runs in several regions and serves traffic from all of them. AWS notes that data corruption still requires backups with a non-zero recovery point.

What does each strategy need from infrastructure?

In backup and restore, everything outside the data is rebuilt at recovery time. AWS says that without infrastructure as code this "may be complex", leading to longer recovery "and possibly exceed your RTO". In pilot light and warm standby, the recovery environment already exists, so the risk moves to drift: AWS advises managing configuration drift in the recovery region so that its infrastructure, data and configuration match what production needs. In active/active, drift and bad changes replicate across regions quickly, so a known-good configuration to roll back to still matters.

Which vendors map to which strategy?

Arpio's Azure launch describes continuous replication that keeps "pilot light" recovery environments in sync with production. Firefly, ControlMonkey and Commvault Cloud Rewind capture configuration so it can be rebuilt, which fits backup and restore and helps keep standby environments in line. See the vendor directory.

What should you take from this lesson?

Pick the strategy per service, from its RTO and RPO. Then check that the infrastructure part of that strategy is captured and tested, not assumed.

Sources

Related