A server room flood, a ransomware attack, a cloud region outage or a mistaken command that deletes a production database: different causes, same result. Your systems are down and the clock is running. A written disaster recovery plan (DR plan) turns that moment from improvisation into a sequence of known steps. This guide walks through writing one, section by section, in a form a small or mid-sized business can realistically maintain.
Disaster recovery vs business continuity
The two terms are related but different. Business continuity covers how the whole organisation keeps operating during a disruption: people, premises, suppliers, communications. Disaster recovery is the IT part of that: restoring systems, applications and data. This article focuses on DR, but your plan should reference whoever owns wider continuity decisions.
Step 1: Define scope and owners
Start with a one-page summary:
- Which systems and locations the plan covers.
- Who can declare a disaster and invoke the plan. Name a primary and a deputy.
- The recovery team, with roles (technical lead, communications, business liaison) and contact details that do not depend on company email being available.
- Where the plan is stored. A copy only on the file server you are trying to recover is no use; keep printed and offsite copies too.
Step 2: List systems and rank them
Make an inventory of every important system and rank it by how much its absence hurts. Talk to department heads, not just IT; they know which outage stops invoicing and which is merely inconvenient. A simple three-tier ranking works well:
- Tier 1: the business cannot trade without it (e-commerce site, ERP, payment processing).
- Tier 2: serious disruption within a day (email, CRM, file shares).
- Tier 3: can wait several days (intranet, archives, reporting).
For each system, record its dependencies. An application server is useless without its database, DNS, authentication service and network.
Step 3: Set RTO and RPO
Two targets drive every technical decision in a disaster recovery plan:
- Recovery Time Objective (RTO): how long the system can be down before the impact becomes unacceptable. For example, 4 hours.
- Recovery Point Objective (RPO): how much data you can afford to lose, measured in time. An RPO of 1 hour means backups or replication must capture changes at least hourly.
| System | Tier | RTO | RPO | Recovery method |
|---|---|---|---|---|
| Online store | 1 | 2 hours | 15 minutes | Warm standby in second region, database replication |
| Accounting | 1 | 8 hours | 24 hours | Restore nightly backup to new VM |
| File server | 2 | 24 hours | 24 hours | Restore from cloud backup |
| Intranet | 3 | 5 days | 1 week | Rebuild from configuration and weekly backup |
Shorter RTOs and RPOs cost more. Be honest about what each system needs rather than marking everything "zero".
Step 4: Choose recovery strategies
Match each system to an approach that meets its targets:
- Backup and restore: cheapest; recovery takes hours to days. Suits most Tier 2 and 3 systems.
- Pilot light: core components such as a replicated database run continuously in a recovery location; application servers are started only when needed.
- Warm standby: a scaled-down but fully working copy runs all the time and is scaled up during a disaster.
- Active-active: two or more sites serve traffic simultaneously; losing one barely affects users. Most expensive and complex.
Cloud platforms make the middle options affordable for smaller businesses, because standby capacity can be small until it is needed. Backups should follow the 3-2-1 rule, with at least one immutable or offline copy to survive ransomware.
Step 5: Write the runbooks
A runbook is a step-by-step recovery procedure for one system, written so that a competent person who did not build the system can follow it under stress. Include:
- Where the backups or replicas are and how to access them, including credentials stored in a password manager or sealed envelope.
- Exact commands and console steps, in order.
- How to verify the system is working afterwards.
- How to redirect users: DNS changes, load balancer updates, VPN settings.
For example, a database restore step might read:
# Restore latest MySQL dump to the recovery server
gunzip < /restore/shop-2024-10-02.sql.gz | mysql -u root -p shop
# Verify row counts against the last known values
mysql -u root -p -e "SELECT COUNT(*) FROM shop.orders;"
Keep DNS TTLs on critical records reasonably short so a failover to a new IP address takes effect quickly. After switching, confirm with a DNS propagation checker.
Step 6: Plan communication
Decide in advance who tells staff, customers and suppliers what, and through which channel. Prepare short message templates for "we are aware of an issue", "service is restored" and, if personal data may be affected, the notifications your legal obligations require. A status page or social media account hosted independently of your main infrastructure is valuable.
Step 7: Test, then test again
An untested plan almost always contains surprises: expired credentials, missing steps, a backup that does not include a key folder. Use escalating test types:
- Tabletop exercise: walk through a scenario around a table, talking through each step. Quick and cheap.
- Component restore: restore one system to an isolated environment and time it against its RTO.
- Full failover: switch real traffic to the recovery environment during a planned window.
Test at least annually, and after major changes. Record the actual recovery times; they are the only evidence your RTOs are achievable.
Step 8: Keep the disaster recovery plan current
Assign an owner to review the plan every six months and whenever systems, suppliers or staff change. Version and date the document so everyone knows they are reading the latest one. If you would like help designing recovery for servers and databases, our server management and database management services include backup and recovery planning.
Key takeaways
- A disaster recovery plan names owners, ranks systems, and sets an RTO and RPO for each.
- Pick recovery strategies per system, from backup and restore to warm standby, according to those targets.
- Write runbooks clear enough for someone else to follow, and store the plan where a disaster cannot reach it.
- Test regularly and update after every significant change.