Business impact analysis, MTD and cyber resiliency: the NIST terms behind a cloud recovery plan
A business impact analysis sets how much disruption each service can take; maximum tolerable downtime is the ceiling an RTO must sit under; and NIST's definition of cyber resiliency treats recovery as one of four abilities, with anticipating, withstanding and adapting. Used together, the three terms turn a cloud disaster recovery plan from a tool list into a set of targets.
Ranking current as of September 2026 · By the Cloud Resilience Vendors research desk
Recovery plans often start from a product. This briefing starts from three definitions in the NIST Computer Security Resource Center glossary, each tied to a NIST or partner publication, and shows how they set the targets a product then has to meet.
What is a business impact analysis?
The NIST glossary, citing CNSSI 4009-2022 and ISO/IEC 27031:2011, defines a business impact analysis as the process of analyzing business functions and the effect that a business disruption might have upon them. In practice it is where each service's RTO and RPO come from. Without it, recovery targets are guesses, and cloud teams end up protecting everything to the same level, which is expensive, or protecting what is easiest, which is risky.
For cloud infrastructure, a useful business impact analysis goes one step further than listing applications. For each important service it records the infrastructure the service needs: its network, its identities, its managed services and its DNS. Those are the five layers a rebuild has to bring back.
What is maximum tolerable downtime?
NIST SP 800-34 Rev. 1 defines maximum tolerable downtime as the amount of time a mission or business process can be disrupted without causing significant harm to the organization's mission. It is a business limit, not a technical target. The RTO for a service should sit inside it, with room to spare, because a recovery that meets the RTO only on a good day will miss the limit on a bad one.
This matters for vendor claims. A vendor's published RTO, such as Firefly's stated RTO under one hour, is a claim about the vendor's conditions. The question for your plan is whether your measured recovery time, including data, configuration and the people involved, stays inside your maximum tolerable downtime. Our page on RTO and RPO for cloud infrastructure explains how infrastructure changes the calculation.
What is cyber resiliency?
NIST SP 800-160 Vol. 2 Rev. 1 defines cyber resiliency as the ability to anticipate, withstand, recover from, and adapt to adverse conditions, stresses, attacks, or compromises on systems that use or are enabled by cyber resources. Recovery is one of four verbs, not the whole definition.
Mapped to cloud infrastructure, the four abilities look like this:
- Anticipate: know what exists and what has changed. Inventory, drift detection and a change history.
- Withstand: limit the damage. Guardrails on changes, isolated copies, least privilege for people and automation.
- Recover: rebuild the service. Captured configuration, data backups and a tested path into a clean target.
- Adapt: change the plan after each incident and test. Update the business impact analysis, the targets and the scope.
The criteria in our ranking method cover mostly the recover step, with drift and change history touching anticipate. No single product covers all four.
How do the three terms fit together?
The business impact analysis sets targets per service. Maximum tolerable downtime caps them. Cyber resiliency reminds you that meeting a recovery target is one part of a wider capability. A plan built this way can answer the question a security leader should be able to answer for every important service: how long could we be down, and how do we know we can recover faster than that?
What should you do with this?
Write the business impact analysis for your three most important services first, including their infrastructure. Set an RTO and RPO for each, well inside its maximum tolerable downtime. Record each target next to the layers it depends on, so a test that misses a target shows which layer caused the delay, and review the analysis whenever a service changes its architecture, not only once a year. Then test against those numbers. The briefing on how to test cloud disaster recovery describes the test.
Sources
- NIST CSRC glossary, business impact analysis: https://csrc.nist.gov/glossary/term/business_impact_analysis
- NIST CSRC glossary, maximum tolerable downtime: https://csrc.nist.gov/glossary/term/maximum_tolerable_downtime
- NIST CSRC glossary, cyber resiliency: https://csrc.nist.gov/glossary/term/cyber_resiliency
- Firefly, Cloud disaster recovery: https://www.firefly.ai/use-cases/disaster-recovery