Blog

Servers & Hosting articles

How to Write an IT Disaster Recovery Plan

How to write an IT disaster recovery plan step by step: set RTO and RPO, map systems, choose recovery options, document runbooks and test them.

4 min read Servers & Hosting

A server room flood, a ransomware attack, a cloud region outage or a mistaken command that deletes a production database: different causes, same result. Your systems are down and the clock is running. A written disaster recovery plan (DR plan) turns that moment from improvisation into a sequence of known steps. This guide walks through writing one, section by section, in a form a small or mid-sized business can realistically maintain.

Disaster recovery vs business continuity

The two terms are related but different. Business continuity covers how the whole organisation keeps operating during a disruption: people, premises, suppliers, communications. Disaster recovery is the IT part of that: restoring systems, applications and data. This article focuses on DR, but your plan should reference whoever owns wider continuity decisions.

Step 1: Define scope and owners

Start with a one-page summary:

  • Which systems and locations the plan covers.
  • Who can declare a disaster and invoke the plan. Name a primary and a deputy.
  • The recovery team, with roles (technical lead, communications, business liaison) and contact details that do not depend on company email being available.
  • Where the plan is stored. A copy only on the file server you are trying to recover is no use; keep printed and offsite copies too.

Step 2: List systems and rank them

Make an inventory of every important system and rank it by how much its absence hurts. Talk to department heads, not just IT; they know which outage stops invoicing and which is merely inconvenient. A simple three-tier ranking works well:

  • Tier 1: the business cannot trade without it (e-commerce site, ERP, payment processing).
  • Tier 2: serious disruption within a day (email, CRM, file shares).
  • Tier 3: can wait several days (intranet, archives, reporting).

For each system, record its dependencies. An application server is useless without its database, DNS, authentication service and network.

Step 3: Set RTO and RPO

Two targets drive every technical decision in a disaster recovery plan:

  • Recovery Time Objective (RTO): how long the system can be down before the impact becomes unacceptable. For example, 4 hours.
  • Recovery Point Objective (RPO): how much data you can afford to lose, measured in time. An RPO of 1 hour means backups or replication must capture changes at least hourly.
SystemTierRTORPORecovery method
Online store12 hours15 minutesWarm standby in second region, database replication
Accounting18 hours24 hoursRestore nightly backup to new VM
File server224 hours24 hoursRestore from cloud backup
Intranet35 days1 weekRebuild from configuration and weekly backup

Shorter RTOs and RPOs cost more. Be honest about what each system needs rather than marking everything "zero".

Step 4: Choose recovery strategies

Match each system to an approach that meets its targets:

  • Backup and restore: cheapest; recovery takes hours to days. Suits most Tier 2 and 3 systems.
  • Pilot light: core components such as a replicated database run continuously in a recovery location; application servers are started only when needed.
  • Warm standby: a scaled-down but fully working copy runs all the time and is scaled up during a disaster.
  • Active-active: two or more sites serve traffic simultaneously; losing one barely affects users. Most expensive and complex.

Cloud platforms make the middle options affordable for smaller businesses, because standby capacity can be small until it is needed. Backups should follow the 3-2-1 rule, with at least one immutable or offline copy to survive ransomware.

Step 5: Write the runbooks

A runbook is a step-by-step recovery procedure for one system, written so that a competent person who did not build the system can follow it under stress. Include:

  1. Where the backups or replicas are and how to access them, including credentials stored in a password manager or sealed envelope.
  2. Exact commands and console steps, in order.
  3. How to verify the system is working afterwards.
  4. How to redirect users: DNS changes, load balancer updates, VPN settings.

For example, a database restore step might read:

# Restore latest MySQL dump to the recovery server
gunzip < /restore/shop-2024-10-02.sql.gz | mysql -u root -p shop
# Verify row counts against the last known values
mysql -u root -p -e "SELECT COUNT(*) FROM shop.orders;"

Keep DNS TTLs on critical records reasonably short so a failover to a new IP address takes effect quickly. After switching, confirm with a DNS propagation checker.

Step 6: Plan communication

Decide in advance who tells staff, customers and suppliers what, and through which channel. Prepare short message templates for "we are aware of an issue", "service is restored" and, if personal data may be affected, the notifications your legal obligations require. A status page or social media account hosted independently of your main infrastructure is valuable.

Step 7: Test, then test again

An untested plan almost always contains surprises: expired credentials, missing steps, a backup that does not include a key folder. Use escalating test types:

  • Tabletop exercise: walk through a scenario around a table, talking through each step. Quick and cheap.
  • Component restore: restore one system to an isolated environment and time it against its RTO.
  • Full failover: switch real traffic to the recovery environment during a planned window.

Test at least annually, and after major changes. Record the actual recovery times; they are the only evidence your RTOs are achievable.

Step 8: Keep the disaster recovery plan current

Assign an owner to review the plan every six months and whenever systems, suppliers or staff change. Version and date the document so everyone knows they are reading the latest one. If you would like help designing recovery for servers and databases, our server management and database management services include backup and recovery planning.

Key takeaways

  • A disaster recovery plan names owners, ranks systems, and sets an RTO and RPO for each.
  • Pick recovery strategies per system, from backup and restore to warm standby, according to those targets.
  • Write runbooks clear enough for someone else to follow, and store the plan where a disaster cannot reach it.
  • Test regularly and update after every significant change.

Need help with this?

Netifi helps businesses around the world with Servers & Hosting. Tell us what you are working on.