ISO/IEC 27001

⌘K
  1. Home
  2. Docs
  3. ISO/IEC 27001
  4. Other Doc
  5. Disaster Recovery Procedure

Disaster Recovery Procedure

1. Purpose

The Disaster Recovery Procedure defines how the organization restores critical information systems, applications, infrastructure, data, and technology services following a disruption or disaster.

The objective is to restore critical services within defined recovery requirements while protecting:

  • Information
  • Systems
  • Customer services
  • Data integrity
  • Security
  • Availability
  • Business operations

Core Principle

Detect → Assess → Activate → Recover → Validate → Stabilize → Communicate → Improve


2. Scope

This procedure applies to recovery of:

  • Production applications
  • Cloud infrastructure
  • Databases
  • Servers
  • Networks
  • Storage
  • SaaS platforms
  • Identity systems
  • Critical integrations
  • CI/CD systems where required
  • Security systems
  • Backup systems
  • Customer-facing services
  • Critical third-party technology dependencies

It covers disasters caused by:

  • Cyberattacks
  • Ransomware
  • Major system failure
  • Cloud outage
  • Data corruption
  • Accidental deletion
  • Infrastructure failure
  • Database failure
  • Network failure
  • Regional outage
  • Critical supplier failure
  • Physical disaster
  • Environmental disruption

3. Disaster Recovery Information

FieldDetails
Procedure ID
Procedure Owner
Technical Owner
Business Owner
Version
Effective Date
Review Date
Approved By
Classification

4. Recovery Objectives

The organization should define recovery requirements for critical services.

Recovery Time Objective — RTO

The maximum targeted time within which a service should be restored following a disruption.

RTO: __________________

Recovery Point Objective — RPO

The maximum acceptable period of data loss measured in time.

RPO: __________________

Maximum Tolerable Downtime

MTD: __________________

Recovery requirements should be based on business impact and risk rather than arbitrary technical targets.


5. Disaster Recovery Roles

RoleResponsibility
Incident ManagerCoordinates overall response
Disaster Recovery LeadCoordinates recovery activities
IT/Cloud TeamRestores infrastructure
Application OwnerValidates application recovery
Database OwnerRestores and validates databases
Security TeamHandles security implications
Business OwnerConfirms business recovery
Communications OwnerCoordinates communications
Supplier OwnerCoordinates third-party recovery
ManagementProvides escalation and decisions

6. Disaster Recovery Activation Criteria

The Disaster Recovery Procedure may be activated when:

☐ Critical production service is unavailable
☐ Recovery cannot be completed through normal incident procedures
☐ Major infrastructure failure occurs
☐ Significant data corruption occurs
☐ Ransomware affects critical systems
☐ Cloud region becomes unavailable
☐ Critical database is lost or corrupted
☐ Major network failure occurs
☐ Critical supplier becomes unavailable
☐ Physical disaster affects technology operations
☐ Business continuity requirements require technology recovery


7. Disaster Severity

Classify the event according to organizational requirements.

SeverityExample
LowNon-critical system disruption
MediumImportant service disruption
HighCritical service significantly affected
CriticalMajor business/customer/service disruption

Severity should determine escalation, recovery priority, communication, and management involvement.


8. Disaster Declaration

The authorized person should determine whether the event qualifies for disaster recovery activation.

Declaration Record

Date/Time: __________________

Declared By: __________________

Reason: __________________

Affected Services: __________________

Initial Severity: __________________

Recovery Lead: __________________


9. Initial Response

Immediately after identifying a major disruption:

  1. Confirm the incident.
  2. Protect personnel where applicable.
  3. Protect remaining systems and information.
  4. Determine affected services.
  5. Prevent further damage.
  6. Activate the appropriate response team.
  7. Assess whether Disaster Recovery should be activated.
  8. Begin incident documentation.

Initial Assessment


10. Safety and Security

During recovery:

☐ Personnel safety considered
☐ Physical access controlled
☐ Compromised systems isolated where required
☐ Evidence preserved
☐ Unauthorized recovery activity prevented
☐ Emergency access controlled
☐ Recovery credentials protected

Recovery activities must not unnecessarily destroy evidence associated with a security incident.


11. Recovery Priorities

Critical systems should be recovered according to business priority.

PrioritySystem/ServiceBusiness ImpactRTORPO
1
2
3

Typical priorities may include:

  1. Identity/authentication
  2. Critical network/connectivity
  3. Core application
  4. Critical database
  5. Customer-facing services
  6. Supporting systems
  7. Non-critical services

The actual order must be based on the organization’s Business Impact Analysis.


12. Dependency Assessment

Before recovery, identify dependencies.

Consider:

  • Identity provider
  • DNS
  • Network
  • Cloud provider
  • Database
  • Storage
  • Encryption keys
  • Secrets
  • Certificates
  • Third-party APIs
  • SaaS services
  • Monitoring
  • Backup systems
  • CI/CD
  • Security systems

Dependency Status

DependencyStatusRecovery Required

13. Backup Verification

Before restoring data:

☐ Backup identified
☐ Backup date verified
☐ Backup integrity checked
☐ Backup source trusted
☐ Backup is appropriate for required RPO
☐ Recovery credentials available
☐ Backup access authorized

Do not restore from a backup suspected of containing malware or corrupted data without appropriate assessment.


14. Recovery Environment

Determine where recovery will occur.

☐ Primary environment
☐ Secondary environment
☐ Alternate cloud region
☐ Disaster recovery environment
☐ Backup environment
☐ Alternate infrastructure
☐ Supplier-provided recovery environment

Recovery Environment


15. Recovery Strategy

Select the appropriate recovery approach.

☐ Restore from backup
☐ Rebuild infrastructure
☐ Failover
☐ Restore database
☐ Restore application
☐ Activate alternate environment
☐ Recreate cloud resources
☐ Restore from Infrastructure-as-Code
☐ Supplier recovery
☐ Manual recovery


16. Infrastructure Recovery

Recover infrastructure in the required sequence.

Typical activities:

  1. Establish network connectivity.
  2. Restore identity and access.
  3. Restore required cloud resources.
  4. Restore security controls.
  5. Restore compute resources.
  6. Restore storage.
  7. Restore databases.
  8. Restore application services.
  9. Restore integrations.
  10. Validate service operation.

17. Cloud Recovery

For cloud environments:

☐ Cloud account available
☐ Administrative access available
☐ MFA available
☐ Recovery region identified
☐ Network configuration available
☐ IAM configuration available
☐ Infrastructure-as-Code available
☐ Security configuration available
☐ Backup available
☐ DNS recovery available
☐ Certificates available
☐ Secrets available through approved mechanism


18. Database Recovery

Database recovery should include:

☐ Recovery point identified
☐ Backup selected
☐ Database restored
☐ Schema validated
☐ Data integrity checked
☐ Access controls restored
☐ Encryption verified
☐ Application connectivity tested
☐ Critical transactions tested

Database Recovery Result


19. Application Recovery

Recover applications using the approved recovery method.

Verify:

☐ Application deployed
☐ Correct version identified
☐ Configuration restored
☐ Secrets available
☐ Database connectivity works
☐ External integrations work
☐ Authentication works
☐ Authorization works
☐ Logging works
☐ Monitoring works


20. Data Integrity Validation

After restoration, validate critical information.

Checks may include:

  • Record counts
  • Database consistency
  • Transaction integrity
  • File integrity
  • Application functionality
  • Customer data
  • Financial transactions
  • Configuration
  • Referential integrity

Validation Result


21. Security Validation

Before returning a recovered system to normal operation:

☐ Security controls active
☐ MFA active
☐ Access controls verified
☐ Privileged accounts reviewed
☐ Network controls active
☐ Encryption enabled
☐ Logging enabled
☐ Monitoring enabled
☐ Vulnerability status reviewed
☐ Malware assessment completed where relevant
☐ Compromised credentials addressed


22. Emergency Access

Emergency access may be required during disaster recovery.

Emergency access must be:

☐ Authorized
☐ Limited
☐ Named
☐ MFA protected where technically possible
☐ Logged
☐ Monitored
☐ Time-bound where practical
☐ Revoked after recovery

All emergency access should be reviewed after the event.


23. Credential Recovery

Recovery may require:

  • Cloud credentials
  • Database credentials
  • Service accounts
  • API keys
  • Certificates
  • Encryption keys
  • Recovery accounts

Credentials must be obtained through approved secure mechanisms.

Do not place credentials in:

  • Recovery documents
  • Email
  • Chat messages
  • Tickets
  • Spreadsheets
  • Source code

24. DNS and Network Recovery

Where applicable:

☐ DNS records restored
☐ Routing restored
☐ Firewall rules restored
☐ Security groups restored
☐ VPN/connectivity restored
☐ Load balancers restored
☐ Network monitoring restored

Validate connectivity before declaring services operational.


25. Third-Party Recovery

For critical suppliers:

☐ Supplier contacted
☐ Supplier incident status obtained
☐ Recovery commitment confirmed
☐ Alternative arrangement assessed
☐ Supplier escalation activated
☐ Customer impact assessed

Critical supplier dependencies should be documented before a disaster occurs.


26. Communication During Recovery

Communication should be coordinated throughout recovery.

Potential stakeholders:

  • Employees
  • Management
  • Customers
  • Suppliers
  • Regulators
  • Security teams
  • Legal
  • Insurance providers
  • Board or leadership

Communication should be:

  • Accurate
  • Authorized
  • Timely
  • Consistent
  • Appropriate to the incident

Do not disclose unverified information.


27. Customer Communication

Where customer services are affected, determine:

☐ Whether customer notification is required
☐ Who approves communication
☐ What information can be disclosed
☐ Expected recovery timeline
☐ Service status
☐ Customer actions required

Customer notification requirements should follow applicable contracts, laws, regulations, and incident procedures.


28. Recovery Testing

Recovery capability should be periodically tested.

Possible tests:

☐ Backup restoration
☐ Database restoration
☐ Application restoration
☐ Cloud failover
☐ Infrastructure rebuild
☐ DNS recovery
☐ Emergency access
☐ Supplier recovery
☐ Full disaster simulation

Testing should reflect the organization’s risk and recovery requirements.


29. Recovery Test Evidence

Record:

FieldDetails
Test Date
Scenario
Systems
RTO
Actual Recovery Time
RPO
Actual Recovery Point
Result
Issues
Corrective Actions

30. Recovery Validation

Before declaring recovery complete:

☐ Critical systems operational
☐ Business functions tested
☐ Data integrity verified
☐ Security controls verified
☐ Access reviewed
☐ Monitoring active
☐ Logging active
☐ Backup restored/confirmed
☐ Dependencies operational
☐ Customer services validated

Recovery Approval

Technical Owner: __________________

Application Owner: __________________

Business Owner: __________________

Date: __________________


31. Return to Normal Operations

After temporary recovery:

  1. Confirm stable operation.
  2. Review temporary controls.
  3. Remove unnecessary emergency access.
  4. Restore standard architecture where appropriate.
  5. Rotate compromised or emergency credentials.
  6. Re-enable normal monitoring.
  7. Confirm backups.
  8. Close temporary infrastructure.
  9. Update documentation.
  10. Complete post-recovery review.

32. Disaster Recovery Closure

The event may be closed when:

☐ Services restored
☐ Business operations validated
☐ Data integrity confirmed
☐ Security risks addressed
☐ Temporary access removed
☐ Temporary infrastructure reviewed
☐ Stakeholders informed
☐ Incident record completed
☐ Recovery evidence retained
☐ Corrective actions recorded


33. Post-Recovery Review

After recovery, evaluate:

  • What happened?
  • What caused the disruption?
  • What worked?
  • What failed?
  • Was the RTO achieved?
  • Was the RPO achieved?
  • Was data lost?
  • Were security controls effective?
  • Were dependencies understood?
  • Were communications effective?
  • Were recovery procedures adequate?

Review Findings


34. Root Cause Analysis

Where appropriate, perform root cause analysis.

Consider:

  • Technical cause
  • Human factors
  • Process weaknesses
  • Configuration
  • Supplier failure
  • Security vulnerability
  • Capacity
  • Change failure
  • Environmental event

Root Cause


35. Corrective Actions

Action IDIssueCorrective ActionOwnerDue DateStatus

Corrective actions should be tracked until implementation and effectiveness are verified.


36. Recovery Evidence

Maintain appropriate evidence such as:

☐ Disaster declaration
☐ Incident record
☐ Recovery decision
☐ Recovery logs
☐ Backup evidence
☐ Restore evidence
☐ Infrastructure recovery evidence
☐ Database recovery evidence
☐ Application validation
☐ Security validation
☐ Communication records
☐ RTO/RPO results
☐ Test results
☐ Corrective actions
☐ Closure approval

Sensitive credentials must not be stored as recovery evidence.


37. Recovery Metrics

Possible metrics include:

MetricResult
Number of DR events
Services affected
Target RTO
Actual recovery time
Target RPO
Actual recovery point
Data loss
Recovery test success rate
Failed recovery tests
Open recovery actions

38. Disaster Recovery Exceptions

Any deviation from the approved recovery strategy should be:

☐ Documented
☐ Risk assessed
☐ Authorized
☐ Recorded
☐ Reviewed after recovery


39. AWS SaaS Startup Example

Environment

A SaaS startup operates its production environment on AWS with:

  • EC2/ECS or Lambda
  • RDS
  • S3
  • CloudFront
  • Route 53
  • IAM
  • CloudWatch
  • Infrastructure-as-Code
  • CI/CD

Scenario

The primary AWS environment experiences a major regional disruption.

Step 1 — Detect

Monitoring identifies that the production service is unavailable.

Step 2 — Assess

The incident team determines:

  • Production is unavailable.
  • Customer access is affected.
  • Primary region is impacted.
  • DR activation criteria are met.

Step 3 — Activate

The Disaster Recovery Lead activates the recovery procedure.

Step 4 — Infrastructure

The team provisions the approved recovery environment using Infrastructure-as-Code.

Step 5 — Database

The latest appropriate RDS backup or replica is restored/activated.

Step 6 — Application

The approved application version is deployed.

Step 7 — Configuration

Approved configuration, secrets, certificates, and security controls are restored.

Step 8 — DNS

Traffic is redirected to the recovered environment.

Step 9 — Validation

The team validates:

  • Login
  • API
  • Database
  • Critical customer workflows
  • Security controls
  • Monitoring
  • Logging

Step 10 — Recovery Confirmation

The business owner confirms that critical customer services are operational.

Step 11 — Post-Recovery

The team:

  • Reviews the incident
  • Confirms data integrity
  • Removes emergency access
  • Reviews temporary infrastructure
  • Documents RTO/RPO performance
  • Creates corrective actions

Audit Trail

Disruption → Detection → Assessment → DR Activation → Infrastructure Recovery → Database Recovery → Application Recovery → DNS/Traffic Recovery → Security Validation → Business Validation → Recovery Approval → Post-Recovery Review


40. Startup-Friendly Disaster Recovery Model

A startup does not necessarily need a large physical disaster recovery facility.

A practical SaaS model may use:

Tier 1 — Critical Systems

  • Production application
  • Primary database
  • Authentication
  • Customer-facing services

Recovery: automated backup/failover where justified.

Tier 2 — Important Systems

  • Monitoring
  • CI/CD
  • Internal applications
  • Supporting infrastructure

Recovery: documented rebuild or restoration.

Tier 3 — Non-Critical Systems

  • Internal productivity tools
  • Non-critical reporting
  • Development systems

Recovery: restore after critical services.

Minimum Recovery Package

A startup should be able to locate:

  • Architecture documentation
  • Asset inventory
  • Critical dependencies
  • Backup information
  • Recovery credentials/process
  • Infrastructure-as-Code
  • Application deployment process
  • DNS information
  • Recovery contacts
  • RTO/RPO requirements
  • Supplier escalation information

41. Common Mistakes

Avoid:

  • Having backups without testing restoration.
  • Assuming cloud automatically means disaster recovery.
  • Not defining RTO/RPO.
  • Keeping backups in the same failure domain without considering the risk.
  • Failing to protect backups from ransomware.
  • Not documenting dependencies.
  • Not testing database recovery.
  • Forgetting DNS and certificates.
  • Forgetting encryption keys or secrets.
  • Giving everyone emergency administrator access.
  • Not documenting emergency access.
  • Declaring recovery based only on infrastructure availability.
  • Failing to validate customer-facing functionality.
  • Never performing recovery tests.
  • Not tracking corrective actions after a failed test.

42. Relationship With Other ISMS Documents

DocumentRelationship
Business Continuity PlanDefines broader business continuity requirements
Business Impact AnalysisIdentifies critical processes and impacts
RTO/RPO AssessmentDefines recovery requirements
ICT Business Continuity PlanAddresses technology continuity
Backup & Restore ProcedureProvides data recovery capability
Incident Response ProcedureHandles the initiating incident
Cloud Security PolicyGoverns cloud recovery
Cloud Exit ChecklistSupports alternate cloud recovery
Access Management ProcedureControls recovery access
Emergency Access ProcedureControls emergency privileges
Emergency Change ProcedureControls emergency recovery changes
Supplier Management ProcedureAddresses critical supplier dependencies
Disaster Recovery Test PlanDefines recovery testing
Disaster Recovery Test ReportRecords test results
Corrective Action TrackerTracks recovery weaknesses

43. ISO/IEC 27001 Connection

Disaster recovery supports the organization’s ability to maintain information security and availability during disruption and to restore systems and information in a controlled manner.

The exact recovery arrangements should be determined based on:

  • ISMS scope
  • Risk assessment
  • Business impact
  • Risk treatment
  • Statement of Applicability
  • Applicable controls
  • Technology architecture
  • Business requirements
  • Legal/regulatory requirements
  • Customer requirements
  • Contractual requirements

This Disaster Recovery Procedure is not itself a universally prescribed ISO/IEC 27001 document. The organization should maintain appropriate documented information and operational evidence to demonstrate that recovery arrangements are planned, implemented, tested, and improved.


44. Audit Evidence Checklist

An auditor should be able to trace:

☐ Business Impact Analysis
☐ Critical systems register
☐ RTO/RPO requirements
☐ Disaster Recovery Plan
☐ Recovery procedure
☐ Backup configuration
☐ Backup monitoring
☐ Restore tests
☐ DR tests
☐ Recovery architecture
☐ Infrastructure-as-Code
☐ Recovery logs
☐ Emergency access records
☐ Recovery validation
☐ RTO/RPO results
☐ Corrective actions
☐ Management review


45. Final Disaster Recovery Audit Trail

For every significant disaster or recovery test, the organization should be able to demonstrate:

What happened?

Which services were affected?

Why was Disaster Recovery activated?

What were the required RTO and RPO?

Which recovery strategy was used?

Which backup or recovery source was used?

Who performed the recovery?

Was data integrity verified?

Were security controls restored?

Was the recovered service validated by the business?

Was the actual recovery time within the required target?

What evidence demonstrates successful recovery?

What corrective actions were identified?

Final Principle

Disaster recovery is not simply restoring servers or databases. It is the controlled restoration of critical services, information, security controls, and business capability within defined recovery requirements, supported by tested procedures and objective evidence.