Hosting Autopsy / Two nameservers, one provider, one bad afternoon

Two nameservers, one provider, one bad afternoon

HOSTING AUTOPSY

6 min read · 1,348 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A small agency of eight people was proud of its redundancy. Its domain listed four nameservers, two on each of two different-looking hostnames, and the owner liked to say that DNS was the one thing they never had to worry about. What nobody at the agency had checked was that all four nameservers belonged to the same DNS company, which had an outage on a Tuesday afternoon.

For about three hours, the agency's website and its email were unreachable for anyone whose resolver did not already have a cached answer. The hosting itself was fine. The server was up, the mailboxes were intact, and the web server log showed a strangely quiet afternoon.

How the afternoon went

The first report came at 14:10, from a client who could not open the agency's site. The account manager tried from her phone and it loaded, which suggested a problem on the client's side. That was a reasonable guess and also wrong: her phone had visited the site that morning and still held the address in its cache.

By 14:40 the pattern was clearer. Some people could reach everything, some could reach nothing, and the split seemed random. Emails sent to the agency were not bouncing at first, then started producing delayed delivery warnings on the sender's side. A client rang to ask whether the agency had closed.

The developer ran a lookup and got no answer at all:

$ dig example.com A

;; ->>HEADER<<- opcode: QUERY, status: SERVFAIL
;; connection timed out; no servers could be reached

Timeouts rather than a clean "no such domain" are the signature of nameservers that are not answering. The domain itself was fine and the records were fine. The servers that were supposed to say so had gone quiet.

Why four nameservers were not enough

When someone types your domain into a browser, their resolver asks the root, then the top-level domain servers, which reply with the list of nameservers you registered. The resolver then asks one of those. If it does not get an answer, it tries the next. That is where redundancy is supposed to come from.

Visitor'sresolver One DNS company, one outage ns1.provider.example ns2.provider.example ns3.provider.example ns4.provider.example shared network, shared control planeno answer from any of them
Four names on the registrar's list look like four independent servers, but a failure at the company behind them takes all four out at once.

The weakness is that the four servers shared a failure domain. They often ran on one network, were managed from one control system and were updated by the same people. A fault in any of those layers affects every server at once. Count of nameservers tells you about hardware failure. It tells you nothing about a company-wide fault, a bad configuration push, an attack on the provider or an expired account.

What cached answers did and did not do

The odd, uneven symptoms came from caching. Every DNS answer carries a time to live, in seconds, that tells resolvers how long they may keep it. The agency's records used a TTL of 3600, an hour, which meant that anyone who had looked up the site in the last hour carried on as normal and anyone else was stranded.

That explains why the first hour felt like a mystery and the second like a catastrophe. As cached entries expired, more and more resolvers went to ask and found nobody home. It also explains why the recovery was ragged: once the provider returned, resolvers that had cached the failure for a few minutes took a while longer to try again.

Email was worse than the website. A sending server that cannot find the recipient's mail records queues the message and retries for hours or days, so most mail was delayed rather than lost. The agency's SPF records live in DNS too, so receivers that could not fetch them sometimes rejected messages outright rather than queueing them.

The wrong turns

For the first forty minutes the team investigated the hosting. They restarted the web server, checked the disk and looked at the error log, all of which showed nothing wrong, because nothing was wrong there. A call to the hosting company was polite and brief: the server was healthy and the queries were not arriving.

The turning point was checking the provider's public status page from a mobile connection, which reported an incident with "degraded name resolution". By then it was nearly 15:00. The outage lasted until a little after 17:00, and the status page said it was caused by a faulty configuration change on their side. The agency had no ability to speed any of it up.

That last point is the real lesson of the afternoon. With every nameserver in one place, there was no switch to flip. Changing nameservers at the registrar would have taken effect only after the registry's own TTL had passed, usually hours to a day, and the new provider would have needed the zone recreated from memory.

The repair

Afterwards the agency added a second, independent DNS provider. There are two usual ways to hold the same records in two places. One is zone transfer: a hidden primary holds the master copy and both providers pull it using AXFR and IXFR. The other is to keep the zone in a tool that publishes to both providers through their APIs. Either works. The point is that there is one source of truth and two places that serve it.

example.com.   86400  IN  NS  ns1.provider-a.example.
example.com.   86400  IN  NS  ns2.provider-a.example.
example.com.   86400  IN  NS  ns1.provider-b.example.
example.com.   86400  IN  NS  ns2.provider-b.example.

With two providers, a resolver that gets no answer from provider A moves on to provider B. The agency also lowered the TTL on its most important records a little, to 300 seconds for the website, so that future changes would take effect faster, and kept a plain-text export of the zone in its password manager.

Try it on your own site

List your nameservers and look at who runs them. The names often give it away, but the addresses are better evidence:

dig NS example.com +short
dig A ns1.provider.example +short
whois 203.0.113.53 | grep -i -E "netname|orgname"

If the four names resolve to addresses in the same network, or the organisation field is identical, you have one provider. Also look at the registrar: if the registrar also hosts your DNS, an account problem could affect both your domain registration and your records. The DNS cheatsheet explains the record types, and the TTL planner helps choose values before a change.

BeforeAfter Provider Ans1 ns2 ns3 ns4single failure domain Provider Ans1 ns2 Provider Bns1 ns2same zone, two independent networks
The number of servers is the same, but in the second layout a single company failing no longer silences the domain.

What the aftermath looked like

The following week the owner asked the practical question: what did this cost? The agency could not put an exact figure on it. Three prospective clients had mentioned the site being down while it was happening, and a retainer client who had been on the point of renewing asked for a call to talk about reliability. Two enquiries that arrived by email during the outage were delayed by several hours, and one of them went to a competitor in the meantime.

The repair itself cost almost nothing. The secondary provider charged a small monthly fee, the migration took an afternoon, and the zone export took ten minutes. The expensive part was the time spent on the wrong theories during the first hour, which a written runbook would have shortened: step one, check the nameservers answer; step two, check the provider status page; step three, only then look at the server. It is a short list, and most hosting incidents answer to a short list.

What would have caught it

PreviousThe robots.txt that told everyone to go awayNextThe wildcard certificate renewal that broke the mail server

More from Hosting Autopsy

Autopsy

The autoscaler that scaled to meet the bots

A startup used auto-scaling so the site would handle spikes. During one weekend a scraper began requesting...

Autopsy

The disk that filled up overnight

A VPS hosted four or five small sites without problems for two years. One morning all of them returned...

Autopsy

The www that was not there

A consultancy launched a redesigned site and shared the address on a printed brochure as...