How to test cloud disaster recovery for infrastructure, not only data
Pick one real application, set an RTO and an RPO for it, and ask each shortlisted vendor to rebuild it into a separate account or region while you time every step. Count the service as recovered only when it serves traffic with its network, identity and DNS in place, not when the data restore finishes.
Ranking current as of September 2026 · By the Cloud Resilience Vendors research desk
Why test the infrastructure and not only the backup?
Most teams already test data restores. Fewer test whether the environment around the data comes back. AWS makes the point in its own guidance on recovery in the cloud: in addition to data, "you must redeploy the infrastructure, configuration, and application code in the recovery Region", and without infrastructure as code that step "will lead to increased recovery times and possibly exceed your RTO". The same guidance warns that "the only error recovery that works is the path you test frequently".
A trial is the cheapest time to run that test. The vendor has an incentive to help, your environment is real, and you have not signed anything yet. The plan below is written for a two to four week trial and works for any product in our ranking, including the data backup platforms.
Step 1: Which application should you test with?
Choose one application that matters to the business but can tolerate a test. It should use more than compute and storage: at least one VPC or virtual network, security groups or network rules, IAM roles, a DNS record, a managed database and, if you can, a SaaS dependency such as your identity provider. An application that only needs a virtual machine and a disk will pass almost any test and tell you little.
Write down every dependency before the trial starts. This list becomes your scoring sheet.
Step 2: What targets should you set before you start?
Set a recovery time objective and a recovery point objective for the application, and set the RPO twice: once for data and once for configuration. A daily configuration capture means up to a day of changes to network rules, IAM policies or DNS can be missing from the restore, even if the database is replicated continuously. Our guide to RTO and RPO for cloud infrastructure covers how to set both.
Decide the recovery target too: the same region, another region, or a clean account the production credentials cannot reach. After ransomware, the clean account is the one that matters.
Step 3: How do you run the rebuild?
Ask the vendor to capture the application, then make a known change (edit a security group rule, add an IAM policy, change a DNS record) and wait for the next capture. Then ask for a rebuild into the target you chose, from the capture taken before your change. You are testing three things at once: coverage, point-in-time accuracy and the rebuild itself.
Vendors do this differently, and the difference shows up in a test:
- Firefly restores by generating Terraform code that your team reviews and applies, rather than changing resources directly. Its documentation notes that restoration availability depends on Terraform provider support, so check the service coverage for every resource on your list.
- Arpio replicates infrastructure and data and runs failover into another region or account on AWS, or another subscription on Azure. Its pricing page lists network sandbox testing in the Premium tier.
- Commvault Cloud Rewind describes point-in-time, in-sync copies of whole applications that it restores without asking you to write infrastructure as code.
- ControlMonkey restores configuration from daily snapshots and generates Terraform, and covers SaaS configuration such as Okta, Entra ID and Cloudflare.
None of these descriptions is a test result. The trial is where you find out.
Step 4: What should you time and record?
Start the clock when you declare the incident, not when you click restore. Record each step with a timestamp:
- Decision and access: how long until someone with the right permissions starts the recovery.
- Infrastructure rebuild: accounts, networks, identity, DNS, managed-service settings.
- Code review and apply, if the tool generates code.
- Data restore, from your backup product or the vendor's.
- Validation: the application serves real requests and the dependency list is complete.
Mark every dependency on your list as restored, restored with manual work, or missing. Note which resources the tool did not support and what your team had to do by hand. The manual items are where a real incident will slow down.
Step 5: What evidence should you keep?
Keep the timestamps, the dependency sheet, the generated code or restore logs, and screenshots of the application working in the recovery target. Security and audit teams will ask for this later, and regulated firms may need it: DORA, which has applied to EU financial entities since 17 January 2025, includes digital operational resilience testing among its core areas. See our briefing on DORA and NIS2 recoverability evidence.
How do you compare vendors after the test?
Put the results side by side: measured time to a working service, dependencies restored without manual work, the configuration RPO you actually observed, and what it would cost to run the same recovery in production. Then read the vendor's published claim next to your number. Firefly states an RTO under one hour and Arpio states RPOs as low as 15 minutes in its Standard tier; both are vendor claims, and your test is the only figure that applies to your environment.
Our buyer's checklist has 24 questions to send before the trial starts, and the ranking shows how each vendor scores on recovery automation and cross-region rebuild from public material.
Sources
- AWS whitepaper, recovery options in the cloud: https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/disaster-recovery-options-in-the-cloud.html
- AWS whitepaper, testing recovery: https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/testing-disaster-recovery.html
- Firefly, backup and recovery documentation: https://docs.firefly.ai/key-features/backup-and-dr.md
- Firefly, cloud disaster recovery use case: https://www.firefly.ai/use-cases/disaster-recovery
- Arpio, Pricing: https://arpio.io/pricing/
- Arpio, Home: https://arpio.io/
- Commvault, Cloud Rewind: https://www.commvault.com/cloud-rewind
- ControlMonkey, Home: https://controlmonkey.io/
- EIOPA, Digital Operational Resilience Act: https://www.eiopa.europa.eu/digital-operational-resilience-act-dora_en