Spare parts do not prevent every outage, because many outages are not caused by parts. Three examples that have been written about publicly make the point.
In March 2021 a fire destroyed one data centre building at OVHcloud's Strasbourg site and damaged another. Customers who had kept backups only in the same location lost data permanently. Redundancy inside a building does not protect against the building.
In October 2021, Facebook and its related services disappeared from the internet for hours after a routine maintenance command withdrew the routes that told the rest of the world how to reach its servers. The same outage also took out internal tools the engineers needed to fix it, which slowed the recovery. The hardware was fine; a configuration change was not.
And in June 2021, a CDN provider called Fastly had a global disruption when a customer's valid configuration change triggered a latent software bug, taking many popular sites offline briefly. Those failures cut across many redundant servers at once because the bug lived in software they all shared.
What redundancy does cover
Redundancy is excellent at one thing: independent, random failures. Disks die, power supplies burn out, fans seize, cables get kicked. If two power supplies fail independently with a small chance each, the chance that both fail in the same hour is the product of two small numbers, which is tiny. That is the arithmetic behind RAID, dual power feeds, bonded network links and clusters of servers, and it works.
The arithmetic only holds when the failures really are independent. Two disks from the same batch, in the same chassis, taking the same vibration and the same heat, are less independent than the brochure assumes. Two servers running the same software with the same configuration pushed at the same minute are not independent at all. The word that matters is correlated: a single cause that hits several supposedly separate parts together.
The usual shapes of a failure that gets past redundancy
| Shape | Example | Why the spare did not help |
|---|---|---|
| Shared location | Fire or flood in the building | Both copies were in the same building |
| Shared software | A bug in code every node runs | The standby had the same bug |
| Shared change | A configuration pushed to everything at once | The mistake arrived on all nodes together |
| Shared dependency | DNS, authentication, a licence server | Everything needed it, nothing could replace it |
| Untested failover | Standby that never took real traffic | It was misconfigured or too small |
| Capacity cascade | One node dies, others overload and follow | The spare could not take the whole load |
Facebook, Fastly and the change that went everywhere
The October 2021 incident is a good lesson in dependencies. The routes in question are announced using BGP; once they were withdrawn, the company's DNS servers, which sat inside the same network, could no longer be reached either, so the names stopped resolving. Staff trying to fix it found that the tools they used to log in and the access control for the buildings relied on the same network. The fix involved people physically going to the sites. The outage was long mostly because the recovery path depended on the thing that was broken.
The Fastly case in June 2021 is the shared-software version. A software update weeks earlier had introduced a bug that sat dormant. A customer pushed a perfectly legitimate configuration setting that triggered it, and a large share of the company's servers returned errors almost at once. The company identified the cause within minutes and most traffic recovered within the hour, but a lot of well-known sites were unreachable meanwhile. No hardware failed in either story.
The cascade: when the spare cannot carry the load
There is one more shape worth a worked example, because it turns a small fault into a large one. Take four web servers sharing traffic, each running at 70 per cent of its capacity at the evening peak (illustrative numbers). One fails. The remaining three must now carry 280 per cent of a server's worth of work between them, which is about 93 per cent each. That sounds survivable, but response times climb steeply near full load, requests start timing out, impatient visitors reload, and the retries add more load. A second server falls over, leaving two with demand well beyond their capacity, and the third and fourth follow within minutes.
The hardware redundancy was real. The capacity planning behind it was not, because "N+1" had been counted in servers rather than in headroom at peak. The fix is dull: run at a lower average utilisation, shed load deliberately (serve a static page, switch off expensive features) rather than collapsing, and rehearse removing one node at the busiest hour you can tolerate.
Untested failover
A failover path that has never carried real traffic is a hope rather than a feature. Common discoveries during the first real test: the standby database is a week behind; the standby has less memory than production; a firewall rule was added to one side only; the certificate on the spare expired last month; the script that promotes the replica needs a password that lives on the dead server.
The good operators handle this by breaking things on purpose. They run game days, switch real traffic to the secondary on a quiet afternoon, pull a power cable in a test room, or inject faults into a live system in small, watched amounts. It is unglamorous work and it finds a surprising amount.
What to do at your own scale
The pattern is that failures often come from changes, shared dependencies and untested failover paths. Good operators schedule drills, stagger rollouts and avoid putting all of a system's control in one place. If you run something important, apply the same thinking at your own scale: keep DNS with a provider separate from your host, keep off-site backups, and test your recovery process before you need it.
Concretely, a reasonable short list for a small business site:
- Registrar login and DNS hosted somewhere other than the web host, with two-factor authentication on both.
- A backup that is stored away from the server and from the provider's account, with a restore tried at least once.
- Credentials for emergencies kept somewhere that does not depend on your own website or mail domain.
- A low TTL on the records you might need to change in a hurry, set before the emergency. The TTL planner shows the trade-offs.
- A written, short list of who to call and which status pages to look at.
Try it on your own site
Start by listing what your site depends on, from the visitor's side. The usual list: a domain name, the registrar, DNS hosting, a server, a database, a certificate, a mail provider, a payment provider. Then ask of each one what happens if it vanishes for six hours. You can see some of it directly:
dig NS example.com +short
dig example.com A +short
openssl s_client -connect example.com:443 -servername example.com 2>/dev/null | openssl x509 -noout -enddate
The first shows which provider hosts your DNS, and whether it is the same company as the web host. The third shows when the certificate expires; the expiry checker does the same thing in a browser. Anything that shows up with all your eggs in one provider's basket is a candidate for a second location.
Quick answers
Does this mean redundancy is a waste?
No. It handles the most frequent failures, which are the ordinary hardware ones. It just does not handle all of them, and those that remain are the ones that make the news.
Which is more likely, a hardware failure or a bad change?
For a mature operator with good hardware redundancy, a change gone wrong is commonly the larger risk. Operators differ, so treat that as a rule of thumb rather than a measurement.
Should I move to a bigger provider to be safe?
Size helps with some risks and creates others, since a large provider has more engineers but also more shared systems. The 2021 stories above involved large, well-resourced companies.
Where can I read the details of past incidents?
Providers often publish post-incident reviews on their status pages, and our own hosting autopsy series goes through composite cases in detail.