1. Purpose
The Disaster Recovery Procedure defines how the organization restores critical information systems, applications, infrastructure, data, and technology services following a disruption or disaster.
The objective is to restore critical services within defined recovery requirements while protecting:
- Information
- Systems
- Customer services
- Data integrity
- Security
- Availability
- Business operations
Core Principle
Detect → Assess → Activate → Recover → Validate → Stabilize → Communicate → Improve
2. Scope
This procedure applies to recovery of:
- Production applications
- Cloud infrastructure
- Databases
- Servers
- Networks
- Storage
- SaaS platforms
- Identity systems
- Critical integrations
- CI/CD systems where required
- Security systems
- Backup systems
- Customer-facing services
- Critical third-party technology dependencies
It covers disasters caused by:
- Cyberattacks
- Ransomware
- Major system failure
- Cloud outage
- Data corruption
- Accidental deletion
- Infrastructure failure
- Database failure
- Network failure
- Regional outage
- Critical supplier failure
- Physical disaster
- Environmental disruption
3. Disaster Recovery Information
| Field | Details |
|---|---|
| Procedure ID | |
| Procedure Owner | |
| Technical Owner | |
| Business Owner | |
| Version | |
| Effective Date | |
| Review Date | |
| Approved By | |
| Classification |
4. Recovery Objectives
The organization should define recovery requirements for critical services.
Recovery Time Objective — RTO
The maximum targeted time within which a service should be restored following a disruption.
RTO: __________________
Recovery Point Objective — RPO
The maximum acceptable period of data loss measured in time.
RPO: __________________
Maximum Tolerable Downtime
MTD: __________________
Recovery requirements should be based on business impact and risk rather than arbitrary technical targets.
5. Disaster Recovery Roles
| Role | Responsibility |
|---|---|
| Incident Manager | Coordinates overall response |
| Disaster Recovery Lead | Coordinates recovery activities |
| IT/Cloud Team | Restores infrastructure |
| Application Owner | Validates application recovery |
| Database Owner | Restores and validates databases |
| Security Team | Handles security implications |
| Business Owner | Confirms business recovery |
| Communications Owner | Coordinates communications |
| Supplier Owner | Coordinates third-party recovery |
| Management | Provides escalation and decisions |
6. Disaster Recovery Activation Criteria
The Disaster Recovery Procedure may be activated when:
☐ Critical production service is unavailable
☐ Recovery cannot be completed through normal incident procedures
☐ Major infrastructure failure occurs
☐ Significant data corruption occurs
☐ Ransomware affects critical systems
☐ Cloud region becomes unavailable
☐ Critical database is lost or corrupted
☐ Major network failure occurs
☐ Critical supplier becomes unavailable
☐ Physical disaster affects technology operations
☐ Business continuity requirements require technology recovery
7. Disaster Severity
Classify the event according to organizational requirements.
| Severity | Example |
|---|---|
| Low | Non-critical system disruption |
| Medium | Important service disruption |
| High | Critical service significantly affected |
| Critical | Major business/customer/service disruption |
Severity should determine escalation, recovery priority, communication, and management involvement.
8. Disaster Declaration
The authorized person should determine whether the event qualifies for disaster recovery activation.
Declaration Record
Date/Time: __________________
Declared By: __________________
Reason: __________________
Affected Services: __________________
Initial Severity: __________________
Recovery Lead: __________________
9. Initial Response
Immediately after identifying a major disruption:
- Confirm the incident.
- Protect personnel where applicable.
- Protect remaining systems and information.
- Determine affected services.
- Prevent further damage.
- Activate the appropriate response team.
- Assess whether Disaster Recovery should be activated.
- Begin incident documentation.
Initial Assessment
10. Safety and Security
During recovery:
☐ Personnel safety considered
☐ Physical access controlled
☐ Compromised systems isolated where required
☐ Evidence preserved
☐ Unauthorized recovery activity prevented
☐ Emergency access controlled
☐ Recovery credentials protected
Recovery activities must not unnecessarily destroy evidence associated with a security incident.
11. Recovery Priorities
Critical systems should be recovered according to business priority.
| Priority | System/Service | Business Impact | RTO | RPO |
|---|---|---|---|---|
| 1 | ||||
| 2 | ||||
| 3 |
Typical priorities may include:
- Identity/authentication
- Critical network/connectivity
- Core application
- Critical database
- Customer-facing services
- Supporting systems
- Non-critical services
The actual order must be based on the organization’s Business Impact Analysis.
12. Dependency Assessment
Before recovery, identify dependencies.
Consider:
- Identity provider
- DNS
- Network
- Cloud provider
- Database
- Storage
- Encryption keys
- Secrets
- Certificates
- Third-party APIs
- SaaS services
- Monitoring
- Backup systems
- CI/CD
- Security systems
Dependency Status
| Dependency | Status | Recovery Required |
|---|---|---|
13. Backup Verification
Before restoring data:
☐ Backup identified
☐ Backup date verified
☐ Backup integrity checked
☐ Backup source trusted
☐ Backup is appropriate for required RPO
☐ Recovery credentials available
☐ Backup access authorized
Do not restore from a backup suspected of containing malware or corrupted data without appropriate assessment.
14. Recovery Environment
Determine where recovery will occur.
☐ Primary environment
☐ Secondary environment
☐ Alternate cloud region
☐ Disaster recovery environment
☐ Backup environment
☐ Alternate infrastructure
☐ Supplier-provided recovery environment
Recovery Environment
15. Recovery Strategy
Select the appropriate recovery approach.
☐ Restore from backup
☐ Rebuild infrastructure
☐ Failover
☐ Restore database
☐ Restore application
☐ Activate alternate environment
☐ Recreate cloud resources
☐ Restore from Infrastructure-as-Code
☐ Supplier recovery
☐ Manual recovery
16. Infrastructure Recovery
Recover infrastructure in the required sequence.
Typical activities:
- Establish network connectivity.
- Restore identity and access.
- Restore required cloud resources.
- Restore security controls.
- Restore compute resources.
- Restore storage.
- Restore databases.
- Restore application services.
- Restore integrations.
- Validate service operation.
17. Cloud Recovery
For cloud environments:
☐ Cloud account available
☐ Administrative access available
☐ MFA available
☐ Recovery region identified
☐ Network configuration available
☐ IAM configuration available
☐ Infrastructure-as-Code available
☐ Security configuration available
☐ Backup available
☐ DNS recovery available
☐ Certificates available
☐ Secrets available through approved mechanism
18. Database Recovery
Database recovery should include:
☐ Recovery point identified
☐ Backup selected
☐ Database restored
☐ Schema validated
☐ Data integrity checked
☐ Access controls restored
☐ Encryption verified
☐ Application connectivity tested
☐ Critical transactions tested
Database Recovery Result
19. Application Recovery
Recover applications using the approved recovery method.
Verify:
☐ Application deployed
☐ Correct version identified
☐ Configuration restored
☐ Secrets available
☐ Database connectivity works
☐ External integrations work
☐ Authentication works
☐ Authorization works
☐ Logging works
☐ Monitoring works
20. Data Integrity Validation
After restoration, validate critical information.
Checks may include:
- Record counts
- Database consistency
- Transaction integrity
- File integrity
- Application functionality
- Customer data
- Financial transactions
- Configuration
- Referential integrity
Validation Result
21. Security Validation
Before returning a recovered system to normal operation:
☐ Security controls active
☐ MFA active
☐ Access controls verified
☐ Privileged accounts reviewed
☐ Network controls active
☐ Encryption enabled
☐ Logging enabled
☐ Monitoring enabled
☐ Vulnerability status reviewed
☐ Malware assessment completed where relevant
☐ Compromised credentials addressed
22. Emergency Access
Emergency access may be required during disaster recovery.
Emergency access must be:
☐ Authorized
☐ Limited
☐ Named
☐ MFA protected where technically possible
☐ Logged
☐ Monitored
☐ Time-bound where practical
☐ Revoked after recovery
All emergency access should be reviewed after the event.
23. Credential Recovery
Recovery may require:
- Cloud credentials
- Database credentials
- Service accounts
- API keys
- Certificates
- Encryption keys
- Recovery accounts
Credentials must be obtained through approved secure mechanisms.
Do not place credentials in:
- Recovery documents
- Chat messages
- Tickets
- Spreadsheets
- Source code
24. DNS and Network Recovery
Where applicable:
☐ DNS records restored
☐ Routing restored
☐ Firewall rules restored
☐ Security groups restored
☐ VPN/connectivity restored
☐ Load balancers restored
☐ Network monitoring restored
Validate connectivity before declaring services operational.
25. Third-Party Recovery
For critical suppliers:
☐ Supplier contacted
☐ Supplier incident status obtained
☐ Recovery commitment confirmed
☐ Alternative arrangement assessed
☐ Supplier escalation activated
☐ Customer impact assessed
Critical supplier dependencies should be documented before a disaster occurs.
26. Communication During Recovery
Communication should be coordinated throughout recovery.
Potential stakeholders:
- Employees
- Management
- Customers
- Suppliers
- Regulators
- Security teams
- Legal
- Insurance providers
- Board or leadership
Communication should be:
- Accurate
- Authorized
- Timely
- Consistent
- Appropriate to the incident
Do not disclose unverified information.
27. Customer Communication
Where customer services are affected, determine:
☐ Whether customer notification is required
☐ Who approves communication
☐ What information can be disclosed
☐ Expected recovery timeline
☐ Service status
☐ Customer actions required
Customer notification requirements should follow applicable contracts, laws, regulations, and incident procedures.
28. Recovery Testing
Recovery capability should be periodically tested.
Possible tests:
☐ Backup restoration
☐ Database restoration
☐ Application restoration
☐ Cloud failover
☐ Infrastructure rebuild
☐ DNS recovery
☐ Emergency access
☐ Supplier recovery
☐ Full disaster simulation
Testing should reflect the organization’s risk and recovery requirements.
29. Recovery Test Evidence
Record:
| Field | Details |
|---|---|
| Test Date | |
| Scenario | |
| Systems | |
| RTO | |
| Actual Recovery Time | |
| RPO | |
| Actual Recovery Point | |
| Result | |
| Issues | |
| Corrective Actions |
30. Recovery Validation
Before declaring recovery complete:
☐ Critical systems operational
☐ Business functions tested
☐ Data integrity verified
☐ Security controls verified
☐ Access reviewed
☐ Monitoring active
☐ Logging active
☐ Backup restored/confirmed
☐ Dependencies operational
☐ Customer services validated
Recovery Approval
Technical Owner: __________________
Application Owner: __________________
Business Owner: __________________
Date: __________________
31. Return to Normal Operations
After temporary recovery:
- Confirm stable operation.
- Review temporary controls.
- Remove unnecessary emergency access.
- Restore standard architecture where appropriate.
- Rotate compromised or emergency credentials.
- Re-enable normal monitoring.
- Confirm backups.
- Close temporary infrastructure.
- Update documentation.
- Complete post-recovery review.
32. Disaster Recovery Closure
The event may be closed when:
☐ Services restored
☐ Business operations validated
☐ Data integrity confirmed
☐ Security risks addressed
☐ Temporary access removed
☐ Temporary infrastructure reviewed
☐ Stakeholders informed
☐ Incident record completed
☐ Recovery evidence retained
☐ Corrective actions recorded
33. Post-Recovery Review
After recovery, evaluate:
- What happened?
- What caused the disruption?
- What worked?
- What failed?
- Was the RTO achieved?
- Was the RPO achieved?
- Was data lost?
- Were security controls effective?
- Were dependencies understood?
- Were communications effective?
- Were recovery procedures adequate?
Review Findings
34. Root Cause Analysis
Where appropriate, perform root cause analysis.
Consider:
- Technical cause
- Human factors
- Process weaknesses
- Configuration
- Supplier failure
- Security vulnerability
- Capacity
- Change failure
- Environmental event
Root Cause
35. Corrective Actions
| Action ID | Issue | Corrective Action | Owner | Due Date | Status |
|---|---|---|---|---|---|
Corrective actions should be tracked until implementation and effectiveness are verified.
36. Recovery Evidence
Maintain appropriate evidence such as:
☐ Disaster declaration
☐ Incident record
☐ Recovery decision
☐ Recovery logs
☐ Backup evidence
☐ Restore evidence
☐ Infrastructure recovery evidence
☐ Database recovery evidence
☐ Application validation
☐ Security validation
☐ Communication records
☐ RTO/RPO results
☐ Test results
☐ Corrective actions
☐ Closure approval
Sensitive credentials must not be stored as recovery evidence.
37. Recovery Metrics
Possible metrics include:
| Metric | Result |
|---|---|
| Number of DR events | |
| Services affected | |
| Target RTO | |
| Actual recovery time | |
| Target RPO | |
| Actual recovery point | |
| Data loss | |
| Recovery test success rate | |
| Failed recovery tests | |
| Open recovery actions |
38. Disaster Recovery Exceptions
Any deviation from the approved recovery strategy should be:
☐ Documented
☐ Risk assessed
☐ Authorized
☐ Recorded
☐ Reviewed after recovery
39. AWS SaaS Startup Example
Environment
A SaaS startup operates its production environment on AWS with:
- EC2/ECS or Lambda
- RDS
- S3
- CloudFront
- Route 53
- IAM
- CloudWatch
- Infrastructure-as-Code
- CI/CD
Scenario
The primary AWS environment experiences a major regional disruption.
Step 1 — Detect
Monitoring identifies that the production service is unavailable.
Step 2 — Assess
The incident team determines:
- Production is unavailable.
- Customer access is affected.
- Primary region is impacted.
- DR activation criteria are met.
Step 3 — Activate
The Disaster Recovery Lead activates the recovery procedure.
Step 4 — Infrastructure
The team provisions the approved recovery environment using Infrastructure-as-Code.
Step 5 — Database
The latest appropriate RDS backup or replica is restored/activated.
Step 6 — Application
The approved application version is deployed.
Step 7 — Configuration
Approved configuration, secrets, certificates, and security controls are restored.
Step 8 — DNS
Traffic is redirected to the recovered environment.
Step 9 — Validation
The team validates:
- Login
- API
- Database
- Critical customer workflows
- Security controls
- Monitoring
- Logging
Step 10 — Recovery Confirmation
The business owner confirms that critical customer services are operational.
Step 11 — Post-Recovery
The team:
- Reviews the incident
- Confirms data integrity
- Removes emergency access
- Reviews temporary infrastructure
- Documents RTO/RPO performance
- Creates corrective actions
Audit Trail
Disruption → Detection → Assessment → DR Activation → Infrastructure Recovery → Database Recovery → Application Recovery → DNS/Traffic Recovery → Security Validation → Business Validation → Recovery Approval → Post-Recovery Review
40. Startup-Friendly Disaster Recovery Model
A startup does not necessarily need a large physical disaster recovery facility.
A practical SaaS model may use:
Tier 1 — Critical Systems
- Production application
- Primary database
- Authentication
- Customer-facing services
Recovery: automated backup/failover where justified.
Tier 2 — Important Systems
- Monitoring
- CI/CD
- Internal applications
- Supporting infrastructure
Recovery: documented rebuild or restoration.
Tier 3 — Non-Critical Systems
- Internal productivity tools
- Non-critical reporting
- Development systems
Recovery: restore after critical services.
Minimum Recovery Package
A startup should be able to locate:
- Architecture documentation
- Asset inventory
- Critical dependencies
- Backup information
- Recovery credentials/process
- Infrastructure-as-Code
- Application deployment process
- DNS information
- Recovery contacts
- RTO/RPO requirements
- Supplier escalation information
41. Common Mistakes
Avoid:
- Having backups without testing restoration.
- Assuming cloud automatically means disaster recovery.
- Not defining RTO/RPO.
- Keeping backups in the same failure domain without considering the risk.
- Failing to protect backups from ransomware.
- Not documenting dependencies.
- Not testing database recovery.
- Forgetting DNS and certificates.
- Forgetting encryption keys or secrets.
- Giving everyone emergency administrator access.
- Not documenting emergency access.
- Declaring recovery based only on infrastructure availability.
- Failing to validate customer-facing functionality.
- Never performing recovery tests.
- Not tracking corrective actions after a failed test.
42. Relationship With Other ISMS Documents
| Document | Relationship |
|---|---|
| Business Continuity Plan | Defines broader business continuity requirements |
| Business Impact Analysis | Identifies critical processes and impacts |
| RTO/RPO Assessment | Defines recovery requirements |
| ICT Business Continuity Plan | Addresses technology continuity |
| Backup & Restore Procedure | Provides data recovery capability |
| Incident Response Procedure | Handles the initiating incident |
| Cloud Security Policy | Governs cloud recovery |
| Cloud Exit Checklist | Supports alternate cloud recovery |
| Access Management Procedure | Controls recovery access |
| Emergency Access Procedure | Controls emergency privileges |
| Emergency Change Procedure | Controls emergency recovery changes |
| Supplier Management Procedure | Addresses critical supplier dependencies |
| Disaster Recovery Test Plan | Defines recovery testing |
| Disaster Recovery Test Report | Records test results |
| Corrective Action Tracker | Tracks recovery weaknesses |
43. ISO/IEC 27001 Connection
Disaster recovery supports the organization’s ability to maintain information security and availability during disruption and to restore systems and information in a controlled manner.
The exact recovery arrangements should be determined based on:
- ISMS scope
- Risk assessment
- Business impact
- Risk treatment
- Statement of Applicability
- Applicable controls
- Technology architecture
- Business requirements
- Legal/regulatory requirements
- Customer requirements
- Contractual requirements
This Disaster Recovery Procedure is not itself a universally prescribed ISO/IEC 27001 document. The organization should maintain appropriate documented information and operational evidence to demonstrate that recovery arrangements are planned, implemented, tested, and improved.
44. Audit Evidence Checklist
An auditor should be able to trace:
☐ Business Impact Analysis
☐ Critical systems register
☐ RTO/RPO requirements
☐ Disaster Recovery Plan
☐ Recovery procedure
☐ Backup configuration
☐ Backup monitoring
☐ Restore tests
☐ DR tests
☐ Recovery architecture
☐ Infrastructure-as-Code
☐ Recovery logs
☐ Emergency access records
☐ Recovery validation
☐ RTO/RPO results
☐ Corrective actions
☐ Management review
45. Final Disaster Recovery Audit Trail
For every significant disaster or recovery test, the organization should be able to demonstrate:
What happened?
Which services were affected?
Why was Disaster Recovery activated?
What were the required RTO and RPO?
Which recovery strategy was used?
Which backup or recovery source was used?
Who performed the recovery?
Was data integrity verified?
Were security controls restored?
Was the recovered service validated by the business?
Was the actual recovery time within the required target?
What evidence demonstrates successful recovery?
What corrective actions were identified?
Final Principle
Disaster recovery is not simply restoring servers or databases. It is the controlled restoration of critical services, information, security controls, and business capability within defined recovery requirements, supported by tested procedures and objective evidence.
