ISO/IEC 27001

⌘K
  1. Home
  2. Docs
  3. ISO/IEC 27001
  4. Other Doc
  5. Disaster Recovery Runbook

Disaster Recovery Runbook

1. Purpose

The Disaster Recovery Runbook provides step-by-step operational instructions for recovering critical ICT services following a major disruption.

Unlike the Disaster Recovery Plan, which defines the overall recovery strategy, this runbook provides the practical execution sequence for the recovery team.

It is intended to help the organization:

  • Recover critical services in a controlled sequence.
  • Restore trusted infrastructure and data.
  • Protect information during recovery.
  • Meet approved RTO/RPO requirements.
  • Coordinate technical recovery activities.
  • Maintain security controls during recovery.
  • Record recovery evidence and decisions.
  • Validate services before returning them to production.
  • Provide a repeatable recovery process during a high-pressure event.

2. Scope

This runbook may be used for:

  • Cloud outage
  • Regional cloud failure
  • Production infrastructure failure
  • Database failure
  • Storage failure
  • Network failure
  • DNS failure
  • Major application failure
  • Ransomware recovery
  • Cloud account compromise
  • Data corruption
  • Backup restoration
  • CI/CD compromise
  • Major supplier technology failure
  • Disaster recovery testing

The runbook should be adapted to the organization’s actual architecture.


3. Runbook Information

FieldDetails
Runbook IDDRR-YYYY-XXXX
Service
EnvironmentProduction / DR / Test
Service Owner
Recovery Lead
Technical Owner
Security Lead
Version
Effective Date
Last Tested
Next Review
Related DR Plan
Related RTO/RPO Assessment
Related BCP

4. Recovery Objectives

ObjectiveTarget
RTO
RPO
Maximum Tolerable Disruption
Recovery Environment
Recovery Region
Critical Data
Critical Dependencies

RTO/RPO values must come from the organization’s approved business continuity and recovery assessments.

They should not be invented by the technical recovery team during an incident unless formally revised and approved.


5. Recovery Scenarios

This runbook should identify the scenario being addressed.

ScenarioApplicable?
Cloud Region Failure☐
Cloud Account Compromise☐
Ransomware☐
Database Failure☐
Data Corruption☐
Application Failure☐
Network Failure☐
DNS Failure☐
Storage Failure☐
CI/CD Compromise☐
Supplier Outage☐
Major Infrastructure Failure☐

Different scenarios may require different recovery procedures.


6. Recovery Team

RoleResponsibility
Incident CommanderOverall incident/recovery coordination
Recovery LeadCoordinates technical recovery
Security LeadSecurity validation and incident containment
Cloud/Infrastructure LeadInfrastructure recovery
Network LeadNetwork/DNS/connectivity recovery
Database LeadDatabase recovery
Application LeadApplication recovery
DevOps LeadCI/CD and deployment
Business OwnerBusiness validation
Supplier OwnerSupplier coordination
Communications LeadStakeholder communication
Executive ManagementMajor business decisions

One person may hold multiple roles in a startup, provided responsibilities remain clear.


7. Recovery Activation Criteria

Activate the runbook when:

  • Normal operational recovery is insufficient.
  • A critical service is unavailable.
  • A major infrastructure failure occurs.
  • Recovery from backup is required.
  • A cloud disaster occurs.
  • A major cyber incident requires trusted-state recovery.
  • The approved DR plan is activated.
  • Management or the Incident Commander authorizes DR activation.

8. Pre-Recovery Safety Check

Do not begin restoration until the recovery team determines whether the original environment is safe to recover from.

Confirm:

  • Incident/disruption recorded.
  • Recovery authority confirmed.
  • Recovery team activated.
  • Affected systems identified.
  • Security incident assessed.
  • Active attacker activity assessed.
  • Recovery environment identified.
  • Backups identified.
  • Recovery point identified.
  • Recovery credentials available.
  • Evidence preservation requirements assessed.
  • Communication channel established.
  • Recovery decision documented.

For ransomware or cloud compromise, do not restore potentially compromised infrastructure without assessing whether the recovery source itself may be affected.


9. Recovery Decision Record

Record:

FieldDetails
Recovery Decision ID
Date/Time
Incident ID
Reason for Recovery
Affected Service
Recovery Strategy
Recovery Point
Recovery Environment
Authorized By
Recovery Lead
Security Review
Expected RTO
Expected RPO

10. Recovery Timeline

Start the recovery clock according to the organization’s approved RTO measurement method.

TimeEventOwnerEvidence
T+00Recovery activatedIncident Commander
T+15Recovery environment preparedCloud Lead
T+30IAM/security validatedSecurity Lead
T+60Network restoredNetwork Lead
T+90Database recovery startedDB Lead
T+150Application deployment startedApp Lead
T+210Application validationApp Lead
T+240Business validationBusiness Owner

The times above are illustrative. Actual recovery times should be measured during the event or test.


11. AWS SaaS Recovery Architecture

For an AWS-based SaaS environment, a typical recovery sequence may be:

AWS Account

↓

IAM / Security

↓

VPC / Network

↓

Security Groups / WAF

↓

S3 / Storage

↓

RDS / Database

↓

ECS / EKS / EC2

↓

Secrets / KMS

↓

Load Balancer

↓

DNS

↓

Monitoring / Logging

↓

Application Validation

↓

Business Validation

The exact sequence depends on the organization’s architecture.


12. Step 1 — Establish Recovery Command

Actions

  1. Activate the recovery team.
  2. Confirm Incident Commander.
  3. Confirm Recovery Lead.
  4. Establish secure communication.
  5. Start recovery decision log.
  6. Record start time.
  7. Confirm target RTO/RPO.
  8. Confirm affected services.
  9. Confirm recovery scenario.

Completion Criteria

  • Recovery authority confirmed.
  • Team assigned.
  • Communication channel active.
  • Recovery objectives confirmed.
  • Recovery record started.

13. Step 2 — Protect Evidence

Where the disruption is security-related:

  • Preserve relevant logs.
  • Preserve incident evidence.
  • Record affected systems.
  • Preserve CloudTrail/security logs.
  • Preserve relevant application/database logs.
  • Preserve configuration information.
  • Record suspicious identities and resources.
  • Maintain evidence traceability.

Recovery should not unnecessarily destroy evidence required for investigation.

For active threats, containment and safety take priority where delay would increase risk.


14. Step 3 — Protect Backups

Before restoration:

  • Identify available backups.
  • Confirm backup date/time.
  • Confirm backup integrity where practical.
  • Identify the last known-good recovery point.
  • Confirm backup is not affected by the incident.
  • Protect backup credentials.
  • Restrict backup deletion.
  • Record selected recovery point.

Recovery Point Record

FieldDetails
Backup ID
Backup Date/Time
Source
Recovery Point
Integrity Check
Selected By
Approved By
Reason Selected

15. Step 4 — Prepare Recovery Environment

Confirm:

  • Recovery account available.
  • Required permissions available.
  • MFA available.
  • Network available.
  • Recovery region/environment available.
  • Infrastructure-as-code available.
  • Required secrets available.
  • Encryption keys available.
  • Recovery tooling available.
  • Monitoring available.

Security Requirements

The recovery environment must use appropriate:

  • MFA
  • Least privilege
  • Secure configuration
  • Network controls
  • Encryption
  • Logging
  • Monitoring
  • Access controls

16. Step 5 — Recover Identity and Access

Restore or validate:

  • Administrative access
  • IAM
  • SSO
  • MFA
  • Recovery roles
  • Service accounts
  • API credentials
  • Secrets
  • Privileged access

Verification

  • MFA functioning
  • Admin roles reviewed
  • Temporary access documented
  • Privileged access restricted
  • Service accounts validated
  • Secrets available
  • Logging enabled
  • Emergency access documented

Do not create permanent privileged access simply to accelerate recovery.


17. Step 6 — Recover Network

Restore:

  • VPC/network
  • Subnets
  • Routing
  • Security groups
  • Network ACLs
  • Firewall
  • WAF
  • Load balancer
  • VPN/private connectivity
  • DNS dependencies

Validation

Confirm:

  • Required connectivity exists.
  • Unnecessary public exposure is not introduced.
  • Security rules match the approved baseline.
  • Monitoring is active.
  • Network logging is enabled where required.

18. Step 7 — Recover Storage

Restore required storage services.

For example:

  • S3
  • File storage
  • Object storage
  • Application configuration
  • Static content

Validate:

  • Correct recovery point.
  • Required data available.
  • Permissions correct.
  • Encryption enabled.
  • Versioning/retention controls available where required.
  • Application access works.

19. Step 8 — Recover Database

Actions

  1. Select approved recovery point.
  2. Restore database.
  3. Apply required configuration.
  4. Validate connectivity.
  5. Validate schema.
  6. Validate critical records.
  7. Validate data integrity.
  8. Enable monitoring.
  9. Restrict administrative access.
  10. Record recovery completion time.

Database Validation

  • Database available.
  • Correct recovery point used.
  • Data integrity checked.
  • Required records available.
  • Application connectivity validated.
  • Security controls enabled.
  • Monitoring active.

20. Step 9 — Recover Application

Restore:

  • Application infrastructure
  • Containers/VMs
  • Application configuration
  • Environment variables
  • Secrets
  • Dependencies
  • APIs
  • Application binaries/images
  • Required certificates

Where possible, deploy from a known-good and trusted source.

Validate:

  • Application starts.
  • Health checks pass.
  • Database connection works.
  • Authentication works.
  • Critical APIs work.
  • Security controls operate.

21. Step 10 — Recover CI/CD

If CI/CD is required for recovery:

  1. Validate source repository.
  2. Confirm repository integrity.
  3. Validate administrator access.
  4. Validate deployment credentials.
  5. Rotate compromised credentials where necessary.
  6. Validate build pipeline.
  7. Validate deployment configuration.
  8. Perform controlled deployment.
  9. Review deployment logs.

For a CI/CD compromise, do not automatically trust the existing pipeline.


22. Step 11 — Recover Secrets and Encryption

Validate:

  • Secrets Manager
  • API keys
  • Service credentials
  • Certificates
  • KMS keys
  • Encryption configuration

If compromise is suspected:

  • Rotate affected credentials.
  • Revoke unnecessary tokens.
  • Replace compromised secrets.
  • Review access.
  • Document rotation.

23. Step 12 — Recover DNS and Traffic

Restore:

  • DNS records
  • Load balancing
  • CDN
  • WAF
  • Routing
  • TLS certificates

Validate:

  • DNS resolution.
  • TLS.
  • Customer connectivity.
  • WAF operation.
  • Application routing.
  • Geographic/network accessibility where applicable.

24. Step 13 — Restore Monitoring and Logging

Before declaring recovery complete, ensure:

  • Cloud logging active.
  • Application logging active.
  • Security monitoring active.
  • Alerting active.
  • Authentication logging active.
  • Infrastructure monitoring active.
  • Database monitoring active.
  • WAF/security alerts active.
  • Backup monitoring active.

A recovered system without monitoring creates a new operational and security risk.


25. Step 14 — Security Validation

Security Lead should validate:

Identity

  • MFA
  • Privileged roles
  • Service accounts
  • Access keys
  • Sessions/tokens

Infrastructure

  • Security groups
  • Network controls
  • Public exposure
  • Encryption
  • Configuration baseline

Application

  • Application security
  • Authentication
  • Authorization
  • Secrets
  • API security

Monitoring

  • Logging
  • Alerts
  • Cloud security monitoring
  • Incident detection

Compromise

If the recovery was caused by a cyber incident:

  • Persistence removed.
  • Unauthorized identities removed.
  • Unauthorized resources removed.
  • Compromised credentials rotated.
  • Malicious configuration removed.
  • Data exposure assessed.

26. Step 15 — Application Validation

The Application Lead should execute agreed critical business tests.

Example:

TestExpected ResultActual ResultStatus
LoginSuccessful
Customer dashboardAvailable
APIOperational
Database transactionSuccessful
File uploadSuccessful
NotificationDelivered
Payment integrationOperational

Only approved critical tests should be used during recovery unless additional testing is authorized.


27. Step 16 — Business Validation

The Business Owner confirms:

  • Critical business process works.
  • Customer service is available.
  • Critical transactions can be completed.
  • Data appears correct.
  • Business dependencies are functioning.
  • Customer impact is understood.

Business Validation Decision

☐ Accepted

☐ Accepted with Restrictions

☐ Not Accepted

Comments:



28. Step 17 — RTO/RPO Measurement

Record actual recovery performance.

MetricTargetActualResult
RTO
RPO
Database Recovery
Application Recovery
Security Validation
Business Validation

If the target was not achieved, record:

  • Reason
  • Business impact
  • Security impact
  • Risk
  • Corrective action
  • Management decision

29. Step 18 — Return to Service

Before production release:

  • Security validation complete.
  • Application validation complete.
  • Data validation complete.
  • Business owner approval obtained.
  • Monitoring active.
  • Logging active.
  • Recovery evidence captured.
  • Communication prepared.
  • Recovery decision recorded.

Then:

Recovery Environment → Controlled Production Service → Enhanced Monitoring


30. Step 19 — Enhanced Monitoring

For an appropriate period following recovery:

  • Monitor authentication.
  • Monitor privileged activity.
  • Monitor network activity.
  • Monitor application errors.
  • Monitor database activity.
  • Monitor security alerts.
  • Monitor customer-impact indicators.
  • Monitor unusual resource activity.

The monitoring period should be based on risk and the nature of the disruption.


31. Step 20 — Remove Temporary Controls

Review and remove:

  • Emergency access
  • Temporary firewall rules
  • Temporary accounts
  • Temporary credentials
  • Temporary infrastructure
  • Temporary DNS settings
  • Temporary security exceptions
  • Temporary monitoring configurations

Every temporary control should have a clear owner and closure decision.


32. Recovery Closure

Recovery may be considered complete when:

  • Critical service restored.
  • RTO/RPO assessed.
  • Data validated.
  • Security validated.
  • Business validated.
  • Monitoring active.
  • Temporary controls removed or formally tracked.
  • Incident status updated.
  • Recovery evidence retained.
  • Outstanding risks identified.
  • Corrective actions assigned.
  • Management notified.

33. Recovery Findings

Record findings such as:

FindingImpactRoot CauseActionOwnerDue Date
Database restore exceeded targetRTO gapRecovery process slowOptimize recovery
Recovery credentials unavailableRecovery delayDocumentation gapUpdate runbook
Monitoring not available initiallySecurity visibility gapRecovery environment incompleteAdd monitoring

34. Recovery Decision Log

Maintain a decision record during recovery.

TimeDecisionReasonDecision MakerEvidence
Activate DRCritical outage
Select backupLatest trusted recovery point
Rotate credentialsCompromise suspected
Restore databaseRequired for service
Return to productionValidation complete

This is particularly valuable during audits and post-incident reviews.


35. Communication Record

Record significant recovery communications.

TimeStakeholderMessageSenderChannel
ManagementDR activated
Technical TeamRecovery started
Customer TeamService impact
ManagementService restored

Communications should contain verified information and avoid speculation.


36. Recovery Evidence Register

Maintain evidence such as:

  • Backup identification
  • Backup integrity checks
  • Recovery timestamps
  • Infrastructure deployment
  • Configuration
  • IAM activity
  • CloudTrail logs
  • Database restoration
  • Application deployment
  • Security validation
  • Monitoring
  • Business validation
  • Communications
  • Decisions
  • Test results
  • Findings
  • Corrective actions

Each important evidence item should have appropriate identification, source, time, ownership, storage, and access controls.


37. DR Runbook Checklist

Activation

  • Incident/disruption recorded
  • DR activation authorized
  • Recovery team activated
  • RTO/RPO confirmed
  • Communication channel established

Assessment

  • Affected services identified
  • Security impact assessed
  • Recovery scenario identified
  • Dependencies identified
  • Recovery point identified

Protection

  • Evidence preserved where required
  • Backups protected
  • Recovery credentials secured
  • Recovery environment protected

Recovery

  • IAM recovered
  • Network recovered
  • Storage recovered
  • Database recovered
  • Application recovered
  • Secrets/KMS validated
  • DNS recovered
  • Monitoring restored
  • Logging restored

Validation

  • Security validation completed
  • Data validation completed
  • Application validation completed
  • Business validation completed
  • RTO measured
  • RPO measured

Closure

  • Service restored
  • Enhanced monitoring enabled
  • Temporary controls reviewed
  • Recovery evidence retained
  • Findings recorded
  • Corrective actions assigned
  • Risk reassessed
  • Management informed
  • Recovery closed

38. Recovery Test Use

This runbook should also be used during controlled DR exercises.

A test should record:

  • Scenario
  • Start time
  • Recovery steps
  • Actual timings
  • Evidence
  • Problems encountered
  • RTO/RPO performance
  • Security validation
  • Business validation
  • Findings
  • Corrective actions
  • Retest requirements

The runbook itself should be updated when testing identifies that the instructions are incomplete or inaccurate.


39. Startup Implementation

A startup does not need a 100-page technical recovery document.

For each critical service, maintain a concise runbook containing:

1. What failed?

2. Who activates recovery?

3. What is the approved RTO/RPO?

4. What is the last trusted recovery point?

5. Where is the recovery environment?

6. What must be recovered first?

7. What commands/configurations/procedures are required?

8. How is security validated?

9. How is the application validated?

10. Who confirms business recovery?

11. What evidence must be retained?

12. What happens after recovery?

The goal is repeatable recovery, not documentation for its own sake.


40. Relationship With Other ISMS Records

The Disaster Recovery Runbook should connect to:

Business Impact Assessment
↓
RTO/RPO Assessment
↓
ICT Dependency Register
↓
ICT Business Continuity Plan
↓
Disaster Recovery Plan
↓
Disaster Recovery Runbook
↓
Backup & Recovery Procedure
↓
DR Test Report
↓
Corrective Action Tracker
↓
Lessons Learned Register
↓
ISMS Improvement Log


41. ISO/IEC 27001 Alignment

The runbook supports the organization’s risk-based arrangements for:

  • ICT readiness for business continuity
  • Information security during disruption
  • Backup
  • Redundancy
  • Access control
  • Privileged access
  • Logging and monitoring
  • Configuration management
  • Change management
  • Incident management
  • Cloud security
  • Supplier continuity
  • Recovery testing

The exact recovery activities should be determined from the organization’s risk assessment, business continuity requirements, technology architecture, and Statement of Applicability.

The runbook is operational evidence that the organization has translated its continuity strategy into executable recovery activities.


42. Audit Evidence

An auditor should be able to select a critical service and trace:

Critical Service
→ Business Impact
→ RTO/RPO
→ Dependencies
→ DR Strategy
→ Runbook
→ Recovery Test/Incident
→ Actual Recovery Time
→ Actual Recovery Point
→ Security Validation
→ Business Validation
→ Findings
→ Corrective Actions
→ Retest
→ Management Review

A documented runbook without evidence that it has been tested provides considerably less assurance than a runbook that has been executed and improved based on actual results.


43. Final Audit Trail

Disruption Detected
→ DR Activation Authorized
→ Recovery Team Activated
→ Impact Assessed
→ Recovery Scenario Identified
→ Evidence Protected
→ Trusted Recovery Point Selected
→ Recovery Environment Prepared
→ Identity/Security Restored
→ Network Restored
→ Data Restored
→ Application Restored
→ Monitoring/Logging Restored
→ Security Validated
→ Application Validated
→ Business Validated
→ RTO/RPO Measured
→ Service Returned to Operation
→ Enhanced Monitoring
→ Temporary Controls Removed
→ Findings Recorded
→ Corrective Actions Assigned
→ Risk Reassessed
→ Recovery Closed


44. Final Principle

A Disaster Recovery Runbook should allow the recovery team to execute recovery under pressure without relying on memory or assumptions.

Activate + Protect + Prepare + Recover + Secure + Validate + Measure + Restore + Monitor + Improve.