Hosting Through the Years / 2016: When DNS fell over

2016: When DNS fell over

HISTORY

4 min read · 821 words

In October 2016, a huge botnet built from poorly secured cameras and routers directed a flood of traffic at Dyn, a major DNS provider. Many well-known sites became unreachable for hours, not because they were down, but because nobody could look up their addresses.

The lesson for site owners was concrete: DNS is a dependency, and relying on a single provider for it is a single point of failure. Secondary DNS from a different company became a mainstream recommendation.

The attack happened on 21 October and came in waves through the day. It was the first time many people outside networking realised that a website can be perfectly healthy on its server and still be out of reach. The server was fine. The signpost was missing.

What was attacked, and with what

The traffic came from Mirai, malware that had spread across internet-connected devices such as IP cameras and home routers. Mirai did not need a clever exploit. It scanned the internet for devices listening for remote logins and tried a short list of factory-default usernames and passwords. A great many devices still had them, and their owners had no idea the camera in the corridor was doing anything at all.

The malware's code had been published a few weeks earlier, which made it easy for others to build their own botnets. Each infected device had little bandwidth by itself. Multiplied by a very large number of them, the total was enough to overwhelm a provider's infrastructure, and what hurt Dyn was a flood of DNS queries aimed at its name servers, which made it hard for ordinary queries to get through.

Why a DNS outage looks like everything being down

When you type a domain into a browser, your device asks a recursive resolver, usually run by your ISP or a public service, to find the address. The resolver works down from the root, then the top-level domain, then asks the domain's own authoritative name servers. Only those last servers know the answer. If they do not respond, the resolver has nothing to give you, and the browser shows an error that says it could not find the server.

BrowserRecursiveresolverRoot.comAuthoritative DNSflooded, no answerYour web server can be healthy and still unreachable if the last hop fails.
The lookup chain ends at the domain's own name servers; if they cannot answer, the rest of the chain has nothing to hand back.

Caching softened the blow for some. A resolver that already has an answer keeps it until the record's TTL runs out. Sites with long TTLs stayed reachable for people whose resolvers had a recent copy, while sites with a TTL of a minute or two lost their answers almost at once. Low TTLs are good for planned migrations and bad for resilience, a trade-off the TTL planner helps you think about.

What the outage taught

The affected sites did not have faulty servers. They shared a provider. That was the finding that stuck: a company could spend heavily on redundant data centres and still have all of its names resolved by one organisation. The more careful operators had listed name servers from two unrelated providers, and when one stopped answering, resolvers moved to the other.

One providerTwo providersns1 and ns2, same companyAttack: all names unreachableProvider AProvider BDownStill answeringIllustrative: an attack on one company leaves the other untouched.
Spreading name servers across two unrelated companies turns a total outage into a degraded one.

Secondary DNS was not a new idea; it had been standard advice since the 1980s, and the old rule was to place secondary servers on separate networks. What the 2016 event did was move it from textbook to checklist. Plenty of domain owners found, in the following weeks, that their "two name servers" were two names for the same provider.

There is a second, smaller lesson about the devices. The botnet was built from products whose owners never logged into them and whose makers never updated them. Changing a default password takes a minute, and disabling remote administration on anything you do not need to reach from outside takes less. After 2016 more vendors began shipping devices with unique passwords, and regulators in several countries started to ask for exactly that.

How to check your own set-up

List your authoritative servers and see who runs them:

dig NS example.com +short
dig NS example.com +trace

If both names sit under the same company's domain, you probably have one provider. A second provider means keeping two copies of the zone in step: either one provider acts as a hidden primary and transfers the zone to the other, or you use a service that pushes the same records to both through their APIs. Common mistakes are forgetting to copy a new record to the second provider and leaving an old NS entry behind. Whichever route you take, update the NS records at your registrar so they list servers from both, and test each one directly with dig @ns1.provider-a.example example.com.

DNS is cheap to make resilient and expensive to lose. Our DNS cheatsheet covers the record types, and the casebook has composite stories about what happens when it goes wrong.
Previous2013 to 2019: Containers, free certificates and HTTPS everywhereNext2018: Privacy law reshapes the registry

More from Hosting Through the Years

History

2000 to 2005: Dot-com aftermath and the rise of blogs

The bursting of the dot-com bubble thinned out the first wave of hosting companies, and the survivors...

History

1998: A new body for names

The domain name system had been run by a small group of individuals and organisations under US government...

History

2004 to 2008: Everyone becomes a publisher

Blogging platforms, photo sharing and social sites made publishing routine. Shared hosting boomed, with...