Hosting Autopsy / The twenty-four hour TTL on the day the server died

The twenty-four hour TTL on the day the server died

HOSTING AUTOPSY

6 min read · 1,342 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A site's DNS records had a TTL of 86,400 seconds, a full day, set years ago by a default. Nobody had chosen it; it was just what the registrar's DNS panel offered when the records were first created, and nobody had a reason to look at it since. One morning the server's disk failed. The team restored the site on a new machine within an hour and updated the A record.

The new server was ready, but for almost a day many visitors were still sent to the dead one by resolvers holding the old answer. Support calls kept coming throughout the day, though the fix had been done by lunchtime.

They lowered the TTL to five minutes on records that matter and raised it only for rarely changed ones. The business here is a composite, a regional estate agent with a property search on its site, but the day it describes is a standard one for anyone who runs DNS for a living.

What a TTL actually does

When a visitor's device wants to reach www.example.com, it asks a recursive resolver, usually run by the visitor's internet provider or a public service. That resolver asks the authoritative servers for the domain, gets an answer, and keeps a copy for the number of seconds the record's TTL (time to live) specifies. Anyone who asks in that period is given the stored answer without the resolver going back to the source.

It is a good system. It makes lookups quick and keeps load off the authoritative servers. But it means that once an answer is out in the world, you cannot recall it. You can change the record at the source in a second. Every resolver that already holds the old answer will carry on using it until its own countdown reaches zero.

The countdown for each resolver started when it last fetched the record, so the expiry times are scattered across the whole TTL window. With a TTL of 86,400, a change made at 08:00 is picked up by some resolvers at 08:01 and by others as late as 08:00 the next morning.

08:00 record changed0 h12 h24 hResolver A: sees new server almost at onceResolver B: old answer until about 12 hResolver C: old answer for the full dayIllustrative: each cache expires on its own schedule
A long TTL spreads the changeover over the whole period, with the unluckiest visitors served the old address for a day.

The day, hour by hour

TimeEvent
06:50Monitoring reports the site down; the disk on the web server has failed
07:15Decision: restore from last night's backup onto a new virtual machine
07:50Site restored, tested using a hosts file entry that points at the new address
08:00A record changed to the new address (203.0.113.40)
08:30Team confirms the site works from the office and from a phone; ticket moved to "resolved"
10:00First wave of calls: "your site is down"
12:30Lunch; the team believes it is finished
AfternoonCalls and emails continue from viewers at various internet providers
Next morningComplaints stop

The team's checks were all honest, and all misleading. The office resolver had a short cache from an earlier test. The phone on mobile data used a different resolver that happened to have expired. Both showed the new site. Neither told them what a customer at a different provider would see.

Why the symptoms were so confusing

Nothing was wrong with the new server, so every direct test passed. The trouble was in a layer the team could not inspect, which is other people's caches. Customers reported the site as down, or sometimes as old, since the dead server occasionally still answered, and the symptoms differed by internet provider, by device and by time of day.

Email was affected too. Mail for the domain was handled by the same machine, with the same long TTL on its records, so some senders kept delivering to the old address. Because the old machine had a failed disk, those messages were not stored anywhere; the senders' servers retried for a while and then returned bounces. A support engineer on the phone to the registrar was told, correctly, that nothing was wrong with the DNS: it had been changed and was being served correctly. The problem was how long the previous answer was allowed to live.

Lowering a TTL after an incident does not help the incident. Resolvers already hold the old record with its old countdown. The TTL must be lowered before the change, by at least as long as the old value.

One practical detail made it worse. The team's status page lived on the same domain and the same server, so it was down too. For most of the morning the only way customers could learn anything was by phoning, which explains why the calls kept coming. A status page on a separate domain and host would have cost almost nothing and saved a great many of those calls.

What they changed afterwards

The team went through the zone and made a decision for each record, rather than applying one value to everything.

RecordOld TTLNew TTLReason
A / AAAA for www and the bare domain86,400300The records you must be able to move in an emergency
MX86,4003,600Mail can be retried, but one hour is a reasonable limit
TXT for SPF and DKIM86,4003,600Rarely changed, moderate caching
NS (delegation)86,40086,400Almost never changes; long is fine

Five minutes means roughly 288 queries a day per busy resolver for each record, which is trivial for any DNS host. The cost is a little extra lookup time for a first visit after expiry, usually milliseconds. Against that, a failover now completes in minutes. The TTL planner can work out how long to lower values before a planned move, and the DNS cheatsheet covers record types.

Planning a change properly

An emergency will always be slower than a planned move, but the same technique helps both. For a planned migration:

  1. At least one old-TTL period before the move (so, a day ahead for 86,400), lower the TTL on the records you will change to 300.
  2. Wait until that old period has fully elapsed, so every cache has fetched the short version.
  3. Make the change and watch it spread over a few minutes.
  4. Leave the old server running for a day anyway, to serve stragglers or at least redirect them.
  5. Raise the TTL again once everything has settled, if you want long caching back.

Some resolvers impose their own minimum or ignore very short values, and browsers keep their own short caches. So even with a perfectly low TTL, expect a few stragglers. The aim is minutes instead of a day.

Worst-case time to follow an A record change, illustrativeTTL 86,40024 hoursTTL 3,6001 hourTTL 3005 minutes
The worst case is the TTL itself, which is why the number you choose on a quiet day decides your next bad day.

Questions that come up

Why not use a TTL of 60 seconds everywhere?

You can, for records that need it. Very low values increase query volume and reduce the benefit of caching, and some resolvers raise them anyway. Five minutes is a good compromise for important records.

Does flushing my own DNS cache fix it?

It fixes your machine and nobody else's, unless the upstream resolver has already expired the record.

Could we have avoided changing DNS at all?

Sometimes. If the new server can take over the old server's IP address, or if the site sits behind a proxy or load balancer with a stable address, you move the backend and DNS stays put. That needs preparing in advance.

What should I do while waiting?

Keep the old address answering if you can, even just with a redirect or a notice, and post a status message. The troubleshooting guide covers the rest.

What would have caught it

PreviousThe domain contact who no longer worked thereNextThe one-line .htaccess edit that took three sites down

More from Hosting Autopsy

Autopsy

The backup that could not be restored

A hobbyist forum had a nightly backup running faithfully for over a year. Every morning the control panel...

Autopsy

The image that cost a month's hosting

A photographer on a pay-by-usage cloud plan published a picture that was picked up by a popular forum. People...

Autopsy

The launch-day database that ran out of connections

A local theatre put tickets for its autumn season on sale at noon and announced it on social media a day...