A site's DNS records had a TTL of 86,400 seconds, a full day, set years ago by a default. Nobody had chosen it; it was just what the registrar's DNS panel offered when the records were first created, and nobody had a reason to look at it since. One morning the server's disk failed. The team restored the site on a new machine within an hour and updated the A record.
The new server was ready, but for almost a day many visitors were still sent to the dead one by resolvers holding the old answer. Support calls kept coming throughout the day, though the fix had been done by lunchtime.
They lowered the TTL to five minutes on records that matter and raised it only for rarely changed ones. The business here is a composite, a regional estate agent with a property search on its site, but the day it describes is a standard one for anyone who runs DNS for a living.
What a TTL actually does
When a visitor's device wants to reach www.example.com, it asks a recursive resolver, usually run by the visitor's internet provider or a public service. That resolver asks the authoritative servers for the domain, gets an answer, and keeps a copy for the number of seconds the record's TTL (time to live) specifies. Anyone who asks in that period is given the stored answer without the resolver going back to the source.
It is a good system. It makes lookups quick and keeps load off the authoritative servers. But it means that once an answer is out in the world, you cannot recall it. You can change the record at the source in a second. Every resolver that already holds the old answer will carry on using it until its own countdown reaches zero.
The countdown for each resolver started when it last fetched the record, so the expiry times are scattered across the whole TTL window. With a TTL of 86,400, a change made at 08:00 is picked up by some resolvers at 08:01 and by others as late as 08:00 the next morning.
The day, hour by hour
| Time | Event |
|---|---|
| 06:50 | Monitoring reports the site down; the disk on the web server has failed |
| 07:15 | Decision: restore from last night's backup onto a new virtual machine |
| 07:50 | Site restored, tested using a hosts file entry that points at the new address |
| 08:00 | A record changed to the new address (203.0.113.40) |
| 08:30 | Team confirms the site works from the office and from a phone; ticket moved to "resolved" |
| 10:00 | First wave of calls: "your site is down" |
| 12:30 | Lunch; the team believes it is finished |
| Afternoon | Calls and emails continue from viewers at various internet providers |
| Next morning | Complaints stop |
The team's checks were all honest, and all misleading. The office resolver had a short cache from an earlier test. The phone on mobile data used a different resolver that happened to have expired. Both showed the new site. Neither told them what a customer at a different provider would see.
Why the symptoms were so confusing
Nothing was wrong with the new server, so every direct test passed. The trouble was in a layer the team could not inspect, which is other people's caches. Customers reported the site as down, or sometimes as old, since the dead server occasionally still answered, and the symptoms differed by internet provider, by device and by time of day.
Email was affected too. Mail for the domain was handled by the same machine, with the same long TTL on its records, so some senders kept delivering to the old address. Because the old machine had a failed disk, those messages were not stored anywhere; the senders' servers retried for a while and then returned bounces. A support engineer on the phone to the registrar was told, correctly, that nothing was wrong with the DNS: it had been changed and was being served correctly. The problem was how long the previous answer was allowed to live.
One practical detail made it worse. The team's status page lived on the same domain and the same server, so it was down too. For most of the morning the only way customers could learn anything was by phoning, which explains why the calls kept coming. A status page on a separate domain and host would have cost almost nothing and saved a great many of those calls.
What they changed afterwards
The team went through the zone and made a decision for each record, rather than applying one value to everything.
| Record | Old TTL | New TTL | Reason |
|---|---|---|---|
| A / AAAA for www and the bare domain | 86,400 | 300 | The records you must be able to move in an emergency |
| MX | 86,400 | 3,600 | Mail can be retried, but one hour is a reasonable limit |
| TXT for SPF and DKIM | 86,400 | 3,600 | Rarely changed, moderate caching |
| NS (delegation) | 86,400 | 86,400 | Almost never changes; long is fine |
Five minutes means roughly 288 queries a day per busy resolver for each record, which is trivial for any DNS host. The cost is a little extra lookup time for a first visit after expiry, usually milliseconds. Against that, a failover now completes in minutes. The TTL planner can work out how long to lower values before a planned move, and the DNS cheatsheet covers record types.
Planning a change properly
An emergency will always be slower than a planned move, but the same technique helps both. For a planned migration:
- At least one old-TTL period before the move (so, a day ahead for 86,400), lower the TTL on the records you will change to 300.
- Wait until that old period has fully elapsed, so every cache has fetched the short version.
- Make the change and watch it spread over a few minutes.
- Leave the old server running for a day anyway, to serve stragglers or at least redirect them.
- Raise the TTL again once everything has settled, if you want long caching back.
Some resolvers impose their own minimum or ignore very short values, and browsers keep their own short caches. So even with a perfectly low TTL, expect a few stragglers. The aim is minutes instead of a day.
Questions that come up
Why not use a TTL of 60 seconds everywhere?
You can, for records that need it. Very low values increase query volume and reduce the benefit of caching, and some resolvers raise them anyway. Five minutes is a good compromise for important records.
Does flushing my own DNS cache fix it?
It fixes your machine and nobody else's, unless the upstream resolver has already expired the record.
Could we have avoided changing DNS at all?
Sometimes. If the new server can take over the old server's IP address, or if the site sits behind a proxy or load balancer with a stable address, you move the backend and DNS stays put. That needs preparing in advance.
What should I do while waiting?
Keep the old address answering if you can, even just with a redirect or a notice, and post a status message. The troubleshooting guide covers the rest.
What would have caught it
- Use a moderate TTL on the records you might need to change in an emergency.
- Know what yours are today.
- Practice a recovery drill and time how long the world takes to follow.
- Test from several resolvers, not only your own, before declaring a fix complete.
- Keep a note of which records need short TTLs and review it whenever the zone changes.