4.2 Business Continuity, Disaster Recovery, and Resilience

Key Takeaways

  • A business impact analysis identifies critical activities and recovery priorities; a business continuity plan is how the business continues; disaster recovery is typically IT restoration — they are not interchangeable labels.
  • Recovery time objective (RTO) and recovery point objective (RPO) are planning criteria that tell you which processes and which backup designs are in scope for the activity under review.
  • Un-tested backups are not evidence of recoverability; ask for the last successful restore of the systems that support the activity.
  • Incident management handles detection, containment, and short-cycle response; disaster recovery is invoked when the disruption exceeds local recovery.
  • Use the BIA's recovery tiers to decide which continuity controls matter for this engagement; do not turn every Part 2 engagement into an enterprise BCP audit or a Part 3 function-management review.
Last updated: August 2026

Business continuity is in CIA Part 2 as a planning recognition topic (A3d), not as a license to write an enterprise resilience manual. When you plan an engagement of an activity, you recognize business resilience, incident management, business impact analysis (BIA), and backup and recovery testing so you can decide which continuity and recovery controls matter for that activity. You are not sitting Part 3 function-management material, and you are not turning every engagement into an audit of the corporate disaster-recovery program.

BIA, BCP, DR, incident management, and resilience are not synonyms

A business impact analysis identifies which activities are critical, what happens as downtime lengthens (lost revenue, safety, regulatory deadlines, customer harm), which systems, people, and vendors they depend on, and how quickly they must return. The BIA produces recovery priorities plus recovery time and recovery point objectives. It is an analysis, not a playbook.

A business continuity plan is how the business keeps operating: manual workarounds, alternate sites, cross-trained staff, crisis communications, and customer commitments during disruption. Disaster recovery is typically the IT restoration plan: rebuild or failover systems, restore data, and reconnect networks. You can have a BCP workaround (paper receiving logs) while DR is still rebuilding the warehouse system. You can also have restored systems (DR success) while the business still cannot invoice because the BCP never assigned who would catch up the backlog.

Business resilience is the broader capability to absorb shocks, adapt, and recover — including cyber events, facility loss, third-party failure, and people loss. Incident management is the operational process for detecting, classifying, containing, communicating, and closing incidents, often on a scale of hours. Disaster recovery is invoked when the incident or event exceeds local recovery: a data-center loss, successful ransomware encryption of production, or a region-wide outage.

Mixing these on the exam is a gift to the item writer. If the stem is about ranking which processes must return first, the tool is the BIA. If it is about who calls customers and how receiving continues on paper, the tool is the BCP. If it is about restoring the general ledger database, the tool is DR. If it is about the ransomware war room in the first four hours, the tool is incident management.

ConceptWhat it answersTypical ownerPlanning use
BIAWhich activities are critical, and how fast must they return?Business with risk/continuitySets recovery priorities and RTO/RPO that drive scope
BCPHow does the business keep operating during disruption?Activity / continuityWorkarounds, staff, communications for this activity
DRHow are systems and data restored?IT / infrastructureRecoverability of the systems this activity depends on
Incident managementHow is an event detected, contained, and closed?IT / operations / cyberPlaybooks for outages and cyber events that hit this activity
ResilienceCan the organization absorb, adapt, and recover?EnterpriseBroader context; do not confuse with one activity's BCP

RTO and RPO as planning criteria

Recovery time objective (RTO) is the maximum acceptable downtime before the impact of an outage becomes intolerable for that activity. Recovery point objective (RPO) is the maximum acceptable data loss, measured in time: an RPO of four hours means backups or replication must be able to restore to a point no older than four hours.

These numbers are planning criteria, not trivia. If payroll's BIA says RTO 24 hours and RPO 4 hours, a weekly backup stored in the same building fails the RPO even if tapes exist, and an untested 72-hour restore runbook fails the RTO. For a treasury payments activity with an RTO of four hours, the engagement risk assessment should treat payment-system recoverability, dual-site processing, and tested failovers as key controls. For a facilities work-order desk with an RTO of 72 hours, the same DR sophistication is not automatically the priority; you would instead ask whether a 72-hour outage is actually acceptable given safety work orders that cannot wait.

The planner compares three things: what the BIA claims, what management's plans and technology can actually deliver, and what the activity under review needs. Mismatches are in-scope risks. An activity labeled Tier 3 in the BIA that in fact feeds a Tier 1 payment process is a dependency gap, not a reason to skip continuity procedures.

Backup existence is not backup recoverability

A3d specifically names backup and recovery testing. Having a backup job in the scheduler is a weak control if restores are never attempted, if backups are not application-consistent, if they sit on the same storage the ransomware encrypts, or if the last successful restore was never documented. Planning questions for the opening interviews:

  • When was the last restore test for the systems that support this activity?
  • Was it a full application restore or a file-level sample?
  • Did the restore meet the RTO, and was the recovered data usable (RPO and integrity)?
  • Who owns offsite, offline, or immutable copies, and have those copies been restored?

If those answers are vague, recoverability is a high residual risk for any engagement of a time-critical activity. The work program then needs procedures that inspect restore-test evidence, not merely a policy that "backups are taken daily." A screenshot of a backup-job success message is not a restore test.

Incident management versus disaster recovery in the work you plan

Incident management and DR intersect, but they are different control systems. Incident management needs severity definitions, on-call rotations, playbooks (including cyber), communication trees, and criteria for escalation to crisis or DR. DR needs restoration runbooks, failover capability, recovery-environment capacity, and tested data restoration. A payment-system outage that is fixed by restarting a service is an incident. The same outage that requires rebuilding the payment database from backup is a DR event.

For planning a treasury or customer-operations engagement, an untested incident playbook is in scope even if "DR" is owned by infrastructure. You are assessing whether the activity under review can detect and contain events that threaten its objectives. You are not assessing whether the chief audit executive has a seat in the enterprise crisis team — that is function-level governance, not this activity's engagement plan.

Which processes are in scope given recovery priorities

Do not attempt to audit the entire enterprise continuity program inside an engagement of one activity. Use the BIA's tiers. If the activity is Tier 1, continuity and recovery controls are key controls and belong in objectives and scope. If the activity is Tier 3, you still ask whether its dependencies (a Tier 1 payment system, a single vendor, a single administrator) create a hidden criticality the BIA missed. If a "noncritical" activity feeds a critical one, that dependency is a planning finding in the making.

Keep the Part 3 boundary clean. How the internal audit function maintains its own continuity, how often the annual plan covers resilience, and how the CAE reports residual risk on enterprise BCP to the board are not the center of this Part 2 objective. The Organizational Resilience Topical Requirement is issued and becomes effective 30 April 2027; scored CIA items on a new Topical Requirement appear at least six months after the effective date. For the current Part 2 exam, A3d still tests the concepts: BIA, BCP, DR, incident management, and backup testing as planning recognition for the activity under review.

Example recovery time objectives used as engagement planning criteria
Test Your Knowledge

During planning, management provides a document that ranks which business activities must return first after a disruption, estimates financial and operational impact by hour of downtime, and lists system and vendor dependencies. Which continuity artifact is this?

A
B
C
D
Test Your Knowledge

Payroll's BIA states an RTO of 24 hours and an RPO of 4 hours. IT shows the planner a daily backup-job log with six months of "success" messages. No restore has been attempted in 18 months, and backups are stored on the same storage array as production. What is the most important planning conclusion?

A
B
C
D
Test Your Knowledge

A treasury engagement is being planned. A payment-system outage last quarter was contained in 90 minutes by restarting a service using the on-call playbook. A separate data-center loss scenario would require rebuilding the payment database from backup at an alternate site. How should the planner treat these two control systems?

A
B
C
D