Learn / How DNS Works, Start to Finish

How DNS Works, Start to Finish

GUIDE

10 min read · 2,206 words

DNS is the system that turns names into addresses and a lot more. This guide follows one lookup from beginning to end, then shows where it goes wrong.

DNS is the system that turns names into addresses and a lot more. It is also the part of the internet that fails in the most confusing way, because when it breaks, everything appears broken at once: the website, the email, the certificate renewal, sometimes the office printer.

This guide follows one lookup from beginning to end, then shows what a zone contains, how delegation works, why changes take time, and how to check each step yourself with a few commands. By the end you should be able to look at a failing domain and say which layer is at fault.

The examples use example.com and addresses from the documentation ranges (203.0.113.x and 2001:db8::). Real output is longer and noisier than the extracts shown here, so treat them as the shape of what you will see.

The cast

Five kinds of participant take part in nearly every lookup.

The distinction between recursive and authoritative is the one that matters most when troubleshooting. A recursive resolver can give a stale answer from its cache. An authoritative server only ever gives its own current data. When two people see different results, the difference nearly always sits in a cache between them and the authoritative server.

A lookup, step by step

Your browser needs the address of www.example.com. It asks the operating system, which asks the resolver. If the resolver has a cached, unexpired answer, it replies immediately. If not, it asks a root server, receives the address of the .com servers, asks them, receives the addresses of example.com's nameservers, and asks one of those. The answer travels back with a TTL, and every layer stores it for that long.

Browser + stub resolver Recursive resolver Root server "ask the .com servers" .com TLD server "ask ns1.example.net" Authoritative server "203.0.113.10, TTL 3600" 1 2 3 4 5. The resolver returns 203.0.113.10 to the stub, and caches it for the TTL. A warm cache skips steps 2 to 4 entirely, which is why most lookups take a millisecond or so.
One uncached lookup: the resolver works down from the root, one referral at a time, and the answer comes back with a TTL.

The resolver does not ask the root for the whole name. It asks for the full name each time but only needs a referral until it reaches a server that knows the answer. In practice it also remembers the TLD servers for a long time, so the root is consulted rarely.

Watching it happen with dig

You can reproduce the walk yourself. dig +trace starts at the root and follows each referral without using your resolver's cache:

dig +trace www.example.com

.                     518400  IN  NS  a.root-servers.net.
...
com.                  172800  IN  NS  a.gtld-servers.net.
...
example.com.          172800  IN  NS  ns1.example.net.
example.com.          172800  IN  NS  ns2.example.net.
...
www.example.com.      3600    IN  A   203.0.113.10

Each block is one hop. Read the first column as the name being delegated, the second as the TTL in seconds, and the last as the thing it points to. The final line is the answer.

To ask a particular server directly, name it with @: dig @ns1.example.net www.example.com asks the authoritative server itself, bypassing every cache. Comparing that answer with dig www.example.com is the quickest way to tell whether a problem is in the zone or in a cache.

Why the transport matters

Ordinary DNS runs over UDP on port 53 because most questions and answers fit in a single small packet. When an answer is too large, for example a long set of DNSSEC signatures, the server tells the client to retry over TCP on the same port. A firewall that allows UDP 53 but drops TCP 53 produces the strangest symptom in the field: small domains work, and certain large ones fail at random. If you run your own nameservers, allow both.

It also explains why a lookup rarely feels slow. A warm cache answers in about a millisecond, and a cold lookup that needs three round trips of roughly 20 to 40 ms each still finishes in a tenth of a second. The pause you notice on a slow page is almost never the name lookup, unless a nameserver has stopped answering and the resolver is waiting out a timeout of a couple of seconds before it tries the next one.

What a zone contains

A zone is the set of records for a domain: address records (A and AAAA), aliases (CNAME), mail routes (MX), text (TXT), nameservers (NS) and administrative data (SOA). See the DNS cheat sheet for each. A small real-world zone looks like this:

example.com.        3600  IN  SOA    ns1.example.net. hostmaster.example.com. 2026100501 7200 3600 1209600 300
example.com.        3600  IN  NS     ns1.example.net.
example.com.        3600  IN  NS     ns2.example.net.
example.com.        3600  IN  A      203.0.113.10
example.com.        3600  IN  AAAA   2001:db8::10
www.example.com.    3600  IN  CNAME  example.com.
example.com.        3600  IN  MX     10 mx1.mail-provider.example.
example.com.        3600  IN  TXT    "v=spf1 mx ~all"

A few rules explain most mistakes. A CNAME cannot sit alongside any other record at the same name, which is why the bare domain (the apex) cannot be a CNAME; many providers offer an ALIAS or flattened record to work around that. An MX record must point to a name, not to an IP address. And the SOA serial number is bumped each time the zone changes, so secondary servers know to fetch the update.

Worked example: one name, three destinations

A village cricket club runs its main site on one host, its fixture app on another, and its mail with a third provider. The zone has to say all of that at once.

example.com.          3600  IN  A      203.0.113.10          ; main site
www.example.com.      3600  IN  CNAME  example.com.         ; same place
fixtures.example.com. 600   IN  A      203.0.113.77          ; app on another server
example.com.          3600  IN  MX     10 mx1.mail-provider.example.
_dmarc.example.com.   3600  IN  TXT    "v=DMARC1; p=none; rua=mailto:[email protected]"
example.com.          3600  IN  CAA    0 issue "letsencrypt.org"

Three things to notice. The fixtures record has a lower TTL because the app moves more often. The MX record names a host in another domain, which is normal; the answer for that host comes from a separate lookup. And the CAA record limits which certificate authorities may issue for the domain, a small control that costs one line.

When a visitor opens fixtures.example.com, the resolver does the walk above again for that name, but it already knows the nameservers for example.com from the earlier lookup, so it skips the root and the TLD and asks the authoritative server straight away. That shortcut is why the second and third lookups for a domain are quicker than the first.

Delegation and glue

The registry for .com stores, for each domain, the names of its nameservers. If those nameservers live inside the domain itself, such as ns1.example.com, the registry also stores their addresses as glue records. Without them, resolvers would be stuck in a loop: to find ns1.example.com they would need to ask ns1.example.com.

This is why changing nameservers is done at the registrar, not in your zone. The NS records inside your zone are a copy; the ones that count are in the parent zone, held by the registry. If the two sets disagree, some resolvers follow one and some the other, and the behaviour is hard to predict. Check both with dig NS example.com and dig +trace.

Lame delegation

A delegation is lame when the registry points at a server that does not actually serve the zone. It is common after moving DNS providers without updating the registrar, or after cancelling the old provider while the delegation still names it. Resolvers try the lame server, get a refusal or silence, and either fail over after a delay or give up. Symptoms are intermittent: slow lookups, random failures.

Caching everywhere

Every layer remembers answers. Your browser does, your operating system does, your router sometimes does, and the resolver always does. That is what makes DNS fast, and what makes changes take time to show. It also caches failures: if a name did not exist when someone asked, that fact may be remembered for a while, usually for the last number in the SOA record (300 seconds in the example above).

record changed Resolver A old answer cached, fully fresh Resolver B old new Resolver C cache was empty: new from the start 0 TTL (1 hour) Illustrative: each resolver switches when its own copy expires, at most one TTL after the change.
Three resolvers, three different moments of switching: nothing happens "everywhere at once".

This is why people see the old site and the new one on different phones. Nothing is stuck; each cache is just counting down. The TTL planner works out when to lower a TTL before a planned move so the countdown is short on the day. Lower it at least one old-TTL ahead, make the change, and raise it again after things settle.

There is a second kind of cache to know about: negative caching. If someone looks up shop.example.com before you create it, the resolver stores "no such name" for a time set by the last field of the SOA record. Create the record a minute later and that person still sees the failure until the negative TTL runs out. When you set up a new subdomain for a launch, create the record a little early and look it up yourself last, not first.

Some software ignores the rules. Certain programming runtimes and old applications cache a resolved address for the whole life of the process, and a few browser extensions or corporate proxies do their own caching. If dig shows the right answer everywhere and one application still reaches the old server, restart the application before suspecting DNS.

Changing DNS safely

A few changes are routine: pointing a name at a new server, adding a mail record, moving DNS to a new provider. The last is the one to plan.

  1. Export or copy every record from the current zone. Include TXT records for mail and verification, which people forget.
  2. Create the same zone at the new provider and compare record by record. Leave nameservers unchanged for now.
  3. Test the new servers directly: dig @ns1.newprovider.example example.com A, and the same for MX and TXT.
  4. Lower the TTL on the NS records and the important names a day beforehand, if the old provider allows it.
  5. Change the nameservers at the registrar. Registries publish the change within minutes to a few hours, but resolvers hold the old delegation for up to its TTL (often 24 to 48 hours for .com).
  6. Leave the old zone running until you are sure no queries still arrive there. Cancelling early is the main cause of lame delegation.
  7. Re-check mail delivery and certificate renewal a day later.
Turn off DNSSEC at the registrar (remove the DS record) before you change DNS provider, unless the new provider can take over the signing keys. A DS record that no longer matches the zone makes validating resolvers return SERVFAIL, and the domain disappears for everyone using them.

Signed and encrypted DNS

Two separate protections sit on top of the basic system, and they are often confused. DNSSEC signs the records in a zone so that a validating resolver can tell whether an answer was altered on its way. It does not hide anything. DNS over HTTPS and DNS over TLS encrypt the conversation between a device and its resolver, so people on the same network cannot read your lookups. They do not prove the answer is true.

For a domain owner, DNSSEC is the one with consequences. It adds a DS record at the registry that must match the keys in your zone. If the keys change and the DS does not, validating resolvers refuse to answer. That is a SERVFAIL for a large share of visitors and none at all for the rest, which makes it very confusing to diagnose. Test with dig +dnssec example.com and look for the ad flag on a validating resolver.

When it breaks

SymptomMeaningFirst check
NXDOMAINThe name does not existTypos, a missing record, an expired domain
SERVFAILThe resolver could not get a valid answerA broken nameserver, or a DNSSEC mistake
TimeoutThe nameservers are not answeringAre they online and reachable on UDP and TCP port 53?
Wrong answerOld data from a cache, or records pointing at the wrong placeAsk the authoritative server directly and compare
REFUSEDThe server will not answer you for that zoneLame delegation or an access rule

The troubleshooting guide covers the wider process, and the glossary defines the terms if any of these are unfamiliar.

Good habits