1. Purpose
The Disaster Recovery Runbook provides step-by-step operational instructions for recovering critical ICT services following a major disruption.
Unlike the Disaster Recovery Plan, which defines the overall recovery strategy, this runbook provides the practical execution sequence for the recovery team.
It is intended to help the organization:
- Recover critical services in a controlled sequence.
- Restore trusted infrastructure and data.
- Protect information during recovery.
- Meet approved RTO/RPO requirements.
- Coordinate technical recovery activities.
- Maintain security controls during recovery.
- Record recovery evidence and decisions.
- Validate services before returning them to production.
- Provide a repeatable recovery process during a high-pressure event.
2. Scope
This runbook may be used for:
- Cloud outage
- Regional cloud failure
- Production infrastructure failure
- Database failure
- Storage failure
- Network failure
- DNS failure
- Major application failure
- Ransomware recovery
- Cloud account compromise
- Data corruption
- Backup restoration
- CI/CD compromise
- Major supplier technology failure
- Disaster recovery testing
The runbook should be adapted to the organization’s actual architecture.
3. Runbook Information
| Field | Details |
|---|---|
| Runbook ID | DRR-YYYY-XXXX |
| Service | |
| Environment | Production / DR / Test |
| Service Owner | |
| Recovery Lead | |
| Technical Owner | |
| Security Lead | |
| Version | |
| Effective Date | |
| Last Tested | |
| Next Review | |
| Related DR Plan | |
| Related RTO/RPO Assessment | |
| Related BCP |
4. Recovery Objectives
| Objective | Target |
|---|---|
| RTO | |
| RPO | |
| Maximum Tolerable Disruption | |
| Recovery Environment | |
| Recovery Region | |
| Critical Data | |
| Critical Dependencies |
RTO/RPO values must come from the organization’s approved business continuity and recovery assessments.
They should not be invented by the technical recovery team during an incident unless formally revised and approved.
5. Recovery Scenarios
This runbook should identify the scenario being addressed.
| Scenario | Applicable? |
|---|---|
| Cloud Region Failure | ☐ |
| Cloud Account Compromise | ☐ |
| Ransomware | ☐ |
| Database Failure | ☐ |
| Data Corruption | ☐ |
| Application Failure | ☐ |
| Network Failure | ☐ |
| DNS Failure | ☐ |
| Storage Failure | ☐ |
| CI/CD Compromise | ☐ |
| Supplier Outage | ☐ |
| Major Infrastructure Failure | ☐ |
Different scenarios may require different recovery procedures.
6. Recovery Team
| Role | Responsibility |
|---|---|
| Incident Commander | Overall incident/recovery coordination |
| Recovery Lead | Coordinates technical recovery |
| Security Lead | Security validation and incident containment |
| Cloud/Infrastructure Lead | Infrastructure recovery |
| Network Lead | Network/DNS/connectivity recovery |
| Database Lead | Database recovery |
| Application Lead | Application recovery |
| DevOps Lead | CI/CD and deployment |
| Business Owner | Business validation |
| Supplier Owner | Supplier coordination |
| Communications Lead | Stakeholder communication |
| Executive Management | Major business decisions |
One person may hold multiple roles in a startup, provided responsibilities remain clear.
7. Recovery Activation Criteria
Activate the runbook when:
- Normal operational recovery is insufficient.
- A critical service is unavailable.
- A major infrastructure failure occurs.
- Recovery from backup is required.
- A cloud disaster occurs.
- A major cyber incident requires trusted-state recovery.
- The approved DR plan is activated.
- Management or the Incident Commander authorizes DR activation.
8. Pre-Recovery Safety Check
Do not begin restoration until the recovery team determines whether the original environment is safe to recover from.
Confirm:
- Incident/disruption recorded.
- Recovery authority confirmed.
- Recovery team activated.
- Affected systems identified.
- Security incident assessed.
- Active attacker activity assessed.
- Recovery environment identified.
- Backups identified.
- Recovery point identified.
- Recovery credentials available.
- Evidence preservation requirements assessed.
- Communication channel established.
- Recovery decision documented.
For ransomware or cloud compromise, do not restore potentially compromised infrastructure without assessing whether the recovery source itself may be affected.
9. Recovery Decision Record
Record:
| Field | Details |
|---|---|
| Recovery Decision ID | |
| Date/Time | |
| Incident ID | |
| Reason for Recovery | |
| Affected Service | |
| Recovery Strategy | |
| Recovery Point | |
| Recovery Environment | |
| Authorized By | |
| Recovery Lead | |
| Security Review | |
| Expected RTO | |
| Expected RPO |
10. Recovery Timeline
Start the recovery clock according to the organization’s approved RTO measurement method.
| Time | Event | Owner | Evidence |
|---|---|---|---|
| T+00 | Recovery activated | Incident Commander | |
| T+15 | Recovery environment prepared | Cloud Lead | |
| T+30 | IAM/security validated | Security Lead | |
| T+60 | Network restored | Network Lead | |
| T+90 | Database recovery started | DB Lead | |
| T+150 | Application deployment started | App Lead | |
| T+210 | Application validation | App Lead | |
| T+240 | Business validation | Business Owner |
The times above are illustrative. Actual recovery times should be measured during the event or test.
11. AWS SaaS Recovery Architecture
For an AWS-based SaaS environment, a typical recovery sequence may be:
AWS Account
↓
IAM / Security
↓
VPC / Network
↓
Security Groups / WAF
↓
S3 / Storage
↓
RDS / Database
↓
ECS / EKS / EC2
↓
Secrets / KMS
↓
Load Balancer
↓
DNS
↓
Monitoring / Logging
↓
Application Validation
↓
Business Validation
The exact sequence depends on the organization’s architecture.
12. Step 1 — Establish Recovery Command
Actions
- Activate the recovery team.
- Confirm Incident Commander.
- Confirm Recovery Lead.
- Establish secure communication.
- Start recovery decision log.
- Record start time.
- Confirm target RTO/RPO.
- Confirm affected services.
- Confirm recovery scenario.
Completion Criteria
- Recovery authority confirmed.
- Team assigned.
- Communication channel active.
- Recovery objectives confirmed.
- Recovery record started.
13. Step 2 — Protect Evidence
Where the disruption is security-related:
- Preserve relevant logs.
- Preserve incident evidence.
- Record affected systems.
- Preserve CloudTrail/security logs.
- Preserve relevant application/database logs.
- Preserve configuration information.
- Record suspicious identities and resources.
- Maintain evidence traceability.
Recovery should not unnecessarily destroy evidence required for investigation.
For active threats, containment and safety take priority where delay would increase risk.
14. Step 3 — Protect Backups
Before restoration:
- Identify available backups.
- Confirm backup date/time.
- Confirm backup integrity where practical.
- Identify the last known-good recovery point.
- Confirm backup is not affected by the incident.
- Protect backup credentials.
- Restrict backup deletion.
- Record selected recovery point.
Recovery Point Record
| Field | Details |
|---|---|
| Backup ID | |
| Backup Date/Time | |
| Source | |
| Recovery Point | |
| Integrity Check | |
| Selected By | |
| Approved By | |
| Reason Selected |
15. Step 4 — Prepare Recovery Environment
Confirm:
- Recovery account available.
- Required permissions available.
- MFA available.
- Network available.
- Recovery region/environment available.
- Infrastructure-as-code available.
- Required secrets available.
- Encryption keys available.
- Recovery tooling available.
- Monitoring available.
Security Requirements
The recovery environment must use appropriate:
- MFA
- Least privilege
- Secure configuration
- Network controls
- Encryption
- Logging
- Monitoring
- Access controls
16. Step 5 — Recover Identity and Access
Restore or validate:
- Administrative access
- IAM
- SSO
- MFA
- Recovery roles
- Service accounts
- API credentials
- Secrets
- Privileged access
Verification
- MFA functioning
- Admin roles reviewed
- Temporary access documented
- Privileged access restricted
- Service accounts validated
- Secrets available
- Logging enabled
- Emergency access documented
Do not create permanent privileged access simply to accelerate recovery.
17. Step 6 — Recover Network
Restore:
- VPC/network
- Subnets
- Routing
- Security groups
- Network ACLs
- Firewall
- WAF
- Load balancer
- VPN/private connectivity
- DNS dependencies
Validation
Confirm:
- Required connectivity exists.
- Unnecessary public exposure is not introduced.
- Security rules match the approved baseline.
- Monitoring is active.
- Network logging is enabled where required.
18. Step 7 — Recover Storage
Restore required storage services.
For example:
- S3
- File storage
- Object storage
- Application configuration
- Static content
Validate:
- Correct recovery point.
- Required data available.
- Permissions correct.
- Encryption enabled.
- Versioning/retention controls available where required.
- Application access works.
19. Step 8 — Recover Database
Actions
- Select approved recovery point.
- Restore database.
- Apply required configuration.
- Validate connectivity.
- Validate schema.
- Validate critical records.
- Validate data integrity.
- Enable monitoring.
- Restrict administrative access.
- Record recovery completion time.
Database Validation
- Database available.
- Correct recovery point used.
- Data integrity checked.
- Required records available.
- Application connectivity validated.
- Security controls enabled.
- Monitoring active.
20. Step 9 — Recover Application
Restore:
- Application infrastructure
- Containers/VMs
- Application configuration
- Environment variables
- Secrets
- Dependencies
- APIs
- Application binaries/images
- Required certificates
Where possible, deploy from a known-good and trusted source.
Validate:
- Application starts.
- Health checks pass.
- Database connection works.
- Authentication works.
- Critical APIs work.
- Security controls operate.
21. Step 10 — Recover CI/CD
If CI/CD is required for recovery:
- Validate source repository.
- Confirm repository integrity.
- Validate administrator access.
- Validate deployment credentials.
- Rotate compromised credentials where necessary.
- Validate build pipeline.
- Validate deployment configuration.
- Perform controlled deployment.
- Review deployment logs.
For a CI/CD compromise, do not automatically trust the existing pipeline.
22. Step 11 — Recover Secrets and Encryption
Validate:
- Secrets Manager
- API keys
- Service credentials
- Certificates
- KMS keys
- Encryption configuration
If compromise is suspected:
- Rotate affected credentials.
- Revoke unnecessary tokens.
- Replace compromised secrets.
- Review access.
- Document rotation.
23. Step 12 — Recover DNS and Traffic
Restore:
- DNS records
- Load balancing
- CDN
- WAF
- Routing
- TLS certificates
Validate:
- DNS resolution.
- TLS.
- Customer connectivity.
- WAF operation.
- Application routing.
- Geographic/network accessibility where applicable.
24. Step 13 — Restore Monitoring and Logging
Before declaring recovery complete, ensure:
- Cloud logging active.
- Application logging active.
- Security monitoring active.
- Alerting active.
- Authentication logging active.
- Infrastructure monitoring active.
- Database monitoring active.
- WAF/security alerts active.
- Backup monitoring active.
A recovered system without monitoring creates a new operational and security risk.
25. Step 14 — Security Validation
Security Lead should validate:
Identity
- MFA
- Privileged roles
- Service accounts
- Access keys
- Sessions/tokens
Infrastructure
- Security groups
- Network controls
- Public exposure
- Encryption
- Configuration baseline
Application
- Application security
- Authentication
- Authorization
- Secrets
- API security
Monitoring
- Logging
- Alerts
- Cloud security monitoring
- Incident detection
Compromise
If the recovery was caused by a cyber incident:
- Persistence removed.
- Unauthorized identities removed.
- Unauthorized resources removed.
- Compromised credentials rotated.
- Malicious configuration removed.
- Data exposure assessed.
26. Step 15 — Application Validation
The Application Lead should execute agreed critical business tests.
Example:
| Test | Expected Result | Actual Result | Status |
|---|---|---|---|
| Login | Successful | ||
| Customer dashboard | Available | ||
| API | Operational | ||
| Database transaction | Successful | ||
| File upload | Successful | ||
| Notification | Delivered | ||
| Payment integration | Operational |
Only approved critical tests should be used during recovery unless additional testing is authorized.
27. Step 16 — Business Validation
The Business Owner confirms:
- Critical business process works.
- Customer service is available.
- Critical transactions can be completed.
- Data appears correct.
- Business dependencies are functioning.
- Customer impact is understood.
Business Validation Decision
☐ Accepted
☐ Accepted with Restrictions
☐ Not Accepted
Comments:
28. Step 17 — RTO/RPO Measurement
Record actual recovery performance.
| Metric | Target | Actual | Result |
|---|---|---|---|
| RTO | |||
| RPO | |||
| Database Recovery | |||
| Application Recovery | |||
| Security Validation | |||
| Business Validation |
If the target was not achieved, record:
- Reason
- Business impact
- Security impact
- Risk
- Corrective action
- Management decision
29. Step 18 — Return to Service
Before production release:
- Security validation complete.
- Application validation complete.
- Data validation complete.
- Business owner approval obtained.
- Monitoring active.
- Logging active.
- Recovery evidence captured.
- Communication prepared.
- Recovery decision recorded.
Then:
Recovery Environment → Controlled Production Service → Enhanced Monitoring
30. Step 19 — Enhanced Monitoring
For an appropriate period following recovery:
- Monitor authentication.
- Monitor privileged activity.
- Monitor network activity.
- Monitor application errors.
- Monitor database activity.
- Monitor security alerts.
- Monitor customer-impact indicators.
- Monitor unusual resource activity.
The monitoring period should be based on risk and the nature of the disruption.
31. Step 20 — Remove Temporary Controls
Review and remove:
- Emergency access
- Temporary firewall rules
- Temporary accounts
- Temporary credentials
- Temporary infrastructure
- Temporary DNS settings
- Temporary security exceptions
- Temporary monitoring configurations
Every temporary control should have a clear owner and closure decision.
32. Recovery Closure
Recovery may be considered complete when:
- Critical service restored.
- RTO/RPO assessed.
- Data validated.
- Security validated.
- Business validated.
- Monitoring active.
- Temporary controls removed or formally tracked.
- Incident status updated.
- Recovery evidence retained.
- Outstanding risks identified.
- Corrective actions assigned.
- Management notified.
33. Recovery Findings
Record findings such as:
| Finding | Impact | Root Cause | Action | Owner | Due Date |
|---|---|---|---|---|---|
| Database restore exceeded target | RTO gap | Recovery process slow | Optimize recovery | ||
| Recovery credentials unavailable | Recovery delay | Documentation gap | Update runbook | ||
| Monitoring not available initially | Security visibility gap | Recovery environment incomplete | Add monitoring |
34. Recovery Decision Log
Maintain a decision record during recovery.
| Time | Decision | Reason | Decision Maker | Evidence |
|---|---|---|---|---|
| Activate DR | Critical outage | |||
| Select backup | Latest trusted recovery point | |||
| Rotate credentials | Compromise suspected | |||
| Restore database | Required for service | |||
| Return to production | Validation complete |
This is particularly valuable during audits and post-incident reviews.
35. Communication Record
Record significant recovery communications.
| Time | Stakeholder | Message | Sender | Channel |
|---|---|---|---|---|
| Management | DR activated | |||
| Technical Team | Recovery started | |||
| Customer Team | Service impact | |||
| Management | Service restored |
Communications should contain verified information and avoid speculation.
36. Recovery Evidence Register
Maintain evidence such as:
- Backup identification
- Backup integrity checks
- Recovery timestamps
- Infrastructure deployment
- Configuration
- IAM activity
- CloudTrail logs
- Database restoration
- Application deployment
- Security validation
- Monitoring
- Business validation
- Communications
- Decisions
- Test results
- Findings
- Corrective actions
Each important evidence item should have appropriate identification, source, time, ownership, storage, and access controls.
37. DR Runbook Checklist
Activation
- Incident/disruption recorded
- DR activation authorized
- Recovery team activated
- RTO/RPO confirmed
- Communication channel established
Assessment
- Affected services identified
- Security impact assessed
- Recovery scenario identified
- Dependencies identified
- Recovery point identified
Protection
- Evidence preserved where required
- Backups protected
- Recovery credentials secured
- Recovery environment protected
Recovery
- IAM recovered
- Network recovered
- Storage recovered
- Database recovered
- Application recovered
- Secrets/KMS validated
- DNS recovered
- Monitoring restored
- Logging restored
Validation
- Security validation completed
- Data validation completed
- Application validation completed
- Business validation completed
- RTO measured
- RPO measured
Closure
- Service restored
- Enhanced monitoring enabled
- Temporary controls reviewed
- Recovery evidence retained
- Findings recorded
- Corrective actions assigned
- Risk reassessed
- Management informed
- Recovery closed
38. Recovery Test Use
This runbook should also be used during controlled DR exercises.
A test should record:
- Scenario
- Start time
- Recovery steps
- Actual timings
- Evidence
- Problems encountered
- RTO/RPO performance
- Security validation
- Business validation
- Findings
- Corrective actions
- Retest requirements
The runbook itself should be updated when testing identifies that the instructions are incomplete or inaccurate.
39. Startup Implementation
A startup does not need a 100-page technical recovery document.
For each critical service, maintain a concise runbook containing:
1. What failed?
2. Who activates recovery?
3. What is the approved RTO/RPO?
4. What is the last trusted recovery point?
5. Where is the recovery environment?
6. What must be recovered first?
7. What commands/configurations/procedures are required?
8. How is security validated?
9. How is the application validated?
10. Who confirms business recovery?
11. What evidence must be retained?
12. What happens after recovery?
The goal is repeatable recovery, not documentation for its own sake.
40. Relationship With Other ISMS Records
The Disaster Recovery Runbook should connect to:
Business Impact Assessment
↓
RTO/RPO Assessment
↓
ICT Dependency Register
↓
ICT Business Continuity Plan
↓
Disaster Recovery Plan
↓
Disaster Recovery Runbook
↓
Backup & Recovery Procedure
↓
DR Test Report
↓
Corrective Action Tracker
↓
Lessons Learned Register
↓
ISMS Improvement Log
41. ISO/IEC 27001 Alignment
The runbook supports the organization’s risk-based arrangements for:
- ICT readiness for business continuity
- Information security during disruption
- Backup
- Redundancy
- Access control
- Privileged access
- Logging and monitoring
- Configuration management
- Change management
- Incident management
- Cloud security
- Supplier continuity
- Recovery testing
The exact recovery activities should be determined from the organization’s risk assessment, business continuity requirements, technology architecture, and Statement of Applicability.
The runbook is operational evidence that the organization has translated its continuity strategy into executable recovery activities.
42. Audit Evidence
An auditor should be able to select a critical service and trace:
Critical Service
→ Business Impact
→ RTO/RPO
→ Dependencies
→ DR Strategy
→ Runbook
→ Recovery Test/Incident
→ Actual Recovery Time
→ Actual Recovery Point
→ Security Validation
→ Business Validation
→ Findings
→ Corrective Actions
→ Retest
→ Management Review
A documented runbook without evidence that it has been tested provides considerably less assurance than a runbook that has been executed and improved based on actual results.
43. Final Audit Trail
Disruption Detected
→ DR Activation Authorized
→ Recovery Team Activated
→ Impact Assessed
→ Recovery Scenario Identified
→ Evidence Protected
→ Trusted Recovery Point Selected
→ Recovery Environment Prepared
→ Identity/Security Restored
→ Network Restored
→ Data Restored
→ Application Restored
→ Monitoring/Logging Restored
→ Security Validated
→ Application Validated
→ Business Validated
→ RTO/RPO Measured
→ Service Returned to Operation
→ Enhanced Monitoring
→ Temporary Controls Removed
→ Findings Recorded
→ Corrective Actions Assigned
→ Risk Reassessed
→ Recovery Closed
44. Final Principle
A Disaster Recovery Runbook should allow the recovery team to execute recovery under pressure without relying on memory or assumptions.
Activate + Protect + Prepare + Recover + Secure + Validate + Measure + Restore + Monitor + Improve.
