Behind the Rack / Disaster recovery sites

Disaster recovery sites

BEHIND THE RACK

4 min read · 770 words

A disaster recovery site is a second location able to take over if the first is lost. It can be a hot standby, running in parallel and ready to take over in seconds; a warm one, ready in minutes or hours; or a cold one, which has to be set up from backups.

Each step toward hot is more expensive. Most small sites do not need one, but they do need backups outside the provider's building. For businesses where an hour of downtime costs real money, a warm standby on a different provider is often the right level.

Two numbers to decide first

Disaster recovery planning comes down to two figures. The recovery time objective (RTO) is how long you can be offline before the damage becomes unacceptable. The recovery point objective (RPO) is how much recent data you can afford to lose. A shop that takes orders every few minutes has a tiny RPO, since losing the last six hours of orders means refunds and angry emails. A brochure site that changes twice a year has an RPO of months and an RTO of a day or two without anyone minding much.

Write both down, with a figure, before shopping for anything. The answers usually turn out to be less demanding than people feared, and that is where the money is saved.

Hot, warm and cold in practice

LevelWhat exists at site twoTypical recoveryData loss
ColdNothing running; backups held off-site, a plan to build from themHours to daysBack to the last backup
WarmA smaller copy of the servers, kept updated every few minutes or hours, not serving visitorsMinutes to hoursMinutes to hours
HotA full copy serving traffic or ready to, with continuous replicationSeconds to minutesSeconds, or none

A cold site on a small budget is just a recent backup stored somewhere else, plus notes on how to rebuild. A warm site for a WordPress shop might be a small VPS at another provider with a copy of the files, a database replica that lags a few minutes behind, and the web server already installed. A hot site means two live copies behind some form of traffic steering, with the awkward problem of keeping writes consistent between them; that is a real engineering project, not a checkbox.

Illustrative: relative cost and relative recovery time Cold cost recovery: hours to days Warm cost minutes to hours Hot cost seconds Bar lengths are only meant to show direction, not real prices.
Moving from cold to hot shortens the outage and lengthens the invoice.

Where the second site should be

A second rack in the same room is not a DR site. Fire, flood, a power failure affecting the whole building or one bad network change can take both. The 2021 fire at the Strasbourg site of a large European provider showed what that looks like for customers whose backups sat next to their servers. A different building in the same city handles most local accidents; a different region handles regional ones; a different provider handles the failure of the company itself, whether it is a billing dispute, a legal action or a bad day for their authentication system. The further apart, the more protection, and the more latency in any replication between them.

A worked example

A small online shop on a VPS takes about 40 orders a day. The owner decides on an RTO of four hours and an RPO of one hour. That rules out a cold-only plan (an untested restore from last night's backup would lose up to a day of orders) and does not justify a hot site. What fits is a warm setup: hourly database dumps and a nightly file sync pushed to a small VPS at a second provider, with the web stack installed and updated monthly. If the primary dies, the owner starts the standby, restores the latest dump, and changes the DNS A record. To make that last step quick, the TTL on the record is set low in advance; the TTL planner helps work out what a low value costs. Cost: a few pounds a month, plus an hour each quarter to rehearse.

Primaryfails Startstandby Restorelatest dump Change DNSA record Visitorson standby Target in this example: under four hours end to end. A low TTL keeps the last step short.
The warm-standby failover in the example, step by step.

Verifying it

The only test that counts is a restore. Once a quarter, bring the standby up on a different address, load the latest backup, and check that a customer could place an order. Time it. Note what was missing: the database user, the cron jobs, the TLS certificate, the mail settings. Compare the result with the RTO you wrote down; if they disagree, one of them has to change. For the DNS side, look at the live TTL with dig example.com A and read the number in the answer section.

PreviousMaintenance windows and change controlNextEnergy, water and the footprint of hosting

More from Behind the Rack

Behind the Rack

Maintenance windows and change control

Many outages begin with a planned change. Serious operators therefore treat changes carefully: peer review...

Behind the Rack

The network

A hosting provider's connection to the internet is not a single cable. It is a set of links to several...

Behind the Rack

Physical security

A facility that stores other people's websites and customer data is serious about who can get in. Typical...