Behind the Rack / Maintenance windows and change control

Maintenance windows and change control

BEHIND THE RACK

4 min read · 779 words

Many outages begin with a planned change. Hardware wears out slowly and predictably, and a provider can buy spares for that. A configuration edit, a firmware update or a routing tweak is different: it takes a system that was working and alters it on purpose, and occasionally the alteration is wrong. Serious operators therefore treat changes carefully: peer review, tests in a staging environment, rollouts to a small group of servers before all of them, and a documented way back.

As a customer you see the result as announced maintenance windows, usually at the quietest hours of the week. If your provider never announces any, it may mean it is very good at rolling changes without downtime, or it may mean it is doing risky work without telling you. Ask which.

What a change actually is

In a data centre, almost anything counts as a change: a new kernel on the web servers, a switch firmware upgrade, a firewall rule, moving a cable, replacing a failed power supply, renewing a certificate, even adding a DNS record on the provider's own zone. The word covers the trivial and the frightening alike, which is why mature teams sort them into classes. A standard change (a routine, pre-approved task with a known result) goes through quickly. A normal change needs review. An emergency change, made while something is already broken, is allowed to skip steps but gets reviewed afterwards.

The classification matters because most of the damage comes from the ones people thought were trivial.

The stages of a careful change

The details vary, but the shape is familiar from nearly every well-run operator.

  1. Write it down: what will change, on which systems, why, and what could go wrong.
  2. Have a second person read it. A reviewer who did not write the plan catches the typo in the subnet mask.
  3. Try it in staging, a smaller copy of the real environment. This finds the obvious failures, though never all of them, since staging rarely matches production exactly.
  4. Roll out to one server, or one rack, and watch it for a while.
  5. Widen to a small percentage, then to everything, pausing between steps.
  6. Keep the old version at hand so that going back is a command, not a rebuild.
Staging test copy One server watch it About 10% pause, compare All servers change complete Roll back if metrics worsen
Each stage widens the blast radius only after the previous one looked healthy; the way back is planned before the first step.

Maintenance windows

Some work cannot be done without interruption: a core router reboot, a power-feed swap, a firmware that needs the whole stack to restart. For these, providers pick a window. The quietest hours are usually early on a weekday morning, local time, though for a provider with customers across the world there is no hour that suits everybody, and a Sunday night for one region is a Monday morning for another.

A decent announcement says what will be done, when it starts and ends, which services are affected, what the expected interruption is (a few seconds of packet loss is not the same as an hour offline), and where updates will appear. Notice of a week is common for planned work with a visible effect, a day or two for routine work, and none for emergencies. The window is a ceiling, not a target: a good team finishes early and says so.

Announce 7 days before Reminder 1 day before Window opens e.g. 03:00 Work done e.g. 03:50 All clear after checks interruption possible here (times illustrative)
A typical sequence of messages around a window; the quiet part in the middle is the only time anything should be affected.

Commands worth running

Look at the provider's status page and its history. A page with a year of entries, each with a start time, an end time and a short post-mortem, tells you more than any claim of 99.99%. Look at how maintenance was described: specific or vague? Check that you are subscribed to notices by email, to an address someone actually reads, because the announcement does no good in the inbox of a person who left in 2022. If your own site matters, add the window to a calendar and look at it afterwards with the uptime calculator to see what a given interruption costs against your target.

Quick answers

Is no announced maintenance a good sign?

Not by itself. Rolling changes across redundant servers can be invisible, but you cannot tell that from silence. Ask how they do it.

Can a provider skip the notice?

For emergency work, yes, and sensibly so: a failing power feed does not wait seven days. The test is whether they explain afterwards.

Should I be doing the same for my own site?

On a small scale, yes. Take a backup, change one thing at a time, look at the result, and keep the previous copy until you are sure.

PreviousThe virtualisation layerNextDisaster recovery sites

More from Behind the Rack

Behind the Rack

Remote hands and hardware swaps

When a disk fails in a server hundreds of kilometres away, somebody has to walk to the rack. Data centre...

Behind the Rack

Why redundancy still fails

Spare parts do not prevent every outage, because many outages are not caused by parts. Three examples that...

Behind the Rack

Power

A data centre is, at heart, a very large and very well-organised electrical installation. Power arrives from...