Hosting Autopsy / The IPv6 record pointing at nothing

The IPv6 record pointing at nothing

HOSTING AUTOPSY

7 min read · 1,432 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A nonprofit that runs a community advice service moved its website to a new host. The migration was done carefully by a volunteer: files copied, database imported, test site checked, then the A record changed to point at the new server. The site came up, the volunteer checked it from the office, and everyone considered the job finished.

What nobody had looked at was the rest of the DNS zone. An old AAAA record stayed in place, pointing at an IPv6 address on the previous server, which no longer served the site.

Visitors on networks that preferred IPv6, which includes many mobile networks, saw timeouts. People on office broadband saw the site normally. Staff checking it from the office could not reproduce the complaints for a week. This is a composite, but a stale AAAA after a move is one of the most reliably confusing faults a support engineer meets, mostly because the person reporting it and the person investigating are on different sides of it.

A and AAAA, briefly

An A record maps a name to an IPv4 address. An AAAA record does the same for IPv6. A site that supports both publishes both, and the visitor's device chooses. Nothing about the website itself says which one to use; it is decided by the client and its network.

The zone for the nonprofit looked like this after the move:

example.org.      300  IN  A     203.0.113.50    ; new server, updated
example.org.      300  IN  AAAA  2001:db8:1::10  ; old server, forgotten
www.example.org.  300  IN  CNAME example.org.

The A record was the one the volunteer knew about, because the hosting company's instructions for the move said "point your A record at the new IP address". Those instructions did not mention IPv6, because the old plan had added it automatically years earlier and nobody had thought of it since.

example.org one name, two records A 203.0.113.50 new server AAAA 2001:db8:1::10 old server Site works Timeout or wrong site
The zone itself is valid. It sends half the world to a server that has stopped serving the site.

Why only some people were affected

When a browser or app connects, it looks up both record types and, if both exist and the device has working IPv6, it tries IPv6 first. Modern clients use a method often called Happy Eyeballs: start with IPv6, and if there is no answer within a fraction of a second, try IPv4 as well. When the IPv6 address is completely unreachable, the visitor may notice only a short delay.

That is not what happened here. The old server was still running, because the previous hosting contract had a month left. It accepted connections on port 443, then either presented the wrong certificate, served a default placeholder, or let the request hang. A connection that succeeds does not trigger a fallback to IPv4. The browser had a working answer from its point of view, and the answer was useless.

The office, like most business broadband at the time, had no IPv6 at all. Every staff device used the A record and saw the new site. Mobile networks commonly provide IPv6 by default, and some are IPv6-only with translation, so phones went to the old address and failed. The split between "works in the office" and "broken on my phone" was the single biggest clue, and it took a week to be recognised.

Office broadband IPv4 only Mobile network IPv6 preferred New server 203.0.113.50 Old server 2001:db8:1::10 uses the A record uses the AAAA record
The same name leads two groups of visitors to two different servers, and the group that could test it was the group that was fine.

How the week went

The first complaint came from a trustee, on her phone at home, who said the site "does not load". The volunteer opened the site on his laptop and it loaded. Told it was probably her signal, she tried again on Wi-Fi and it worked, which seemed to confirm that.

Over the next few days came more reports, each easy to dismiss alone: a donor on a train, a user on a tablet using mobile data, a person abroad. The shared feature was never obvious because the reporters didn't know it themselves. Page-view statistics showed a gradual dip, mostly in mobile traffic, which was put down to the school holidays.

The turning point was a support engineer at the new host. She was asked why the site was "slow for some people" and ran a lookup for both record types as a matter of routine.

Diagnosis

Two commands showed it. Neither required any special access.

$ dig +short A example.org
203.0.113.50
$ dig +short AAAA example.org
2001:db8:1::10

The second answer should have been empty. To confirm the old address was the problem, she forced each protocol in turn:

$ curl -4 -sI https://example.org/ | head -1
HTTP/2 200
$ curl -6 -sI --max-time 10 https://example.org/
curl: (28) Operation timed out after 10001 milliseconds

That was the whole case. The site worked over IPv4 and failed over IPv6. The same check from the office, which has no IPv6, would never have shown it, which is why testing from a second kind of network matters.

The fix

The decision was between deleting the AAAA record and pointing it at the new server. The new host did offer IPv6, so either was possible. The nonprofit chose deletion for now: it was the quick change, and they could add a correct record later after verifying that the new server answered properly on IPv6.

The record had a five-minute TTL, so most resolvers dropped it quickly. A few caches, including some on mobile carriers, held it for longer than the stated TTL, and the last complaints dried up after about a day. A lower TTL set before the move would have made this faster. The TTL planner shows how to schedule that, and the DNS cheat sheet covers record types in general.

The other records worth auditing

The AAAA record was the one that bit, but a migration leaves other loose ends. A short audit of the whole zone, done once with the old and new hosting details side by side, catches most of them.

RecordWhat to check after a move
A and AAAABoth point at the new host, or the AAAA is absent
CNAME (www, shop, blog)Target still exists and is where you expect
MXUnchanged unless mail moved too
TXT (SPF, verification)SPF names the new sending servers, old includes removed
Subdomains such as staging or mailNot pointing at a server you have cancelled

The nonprofit's zone turned out to have a forgotten staging subdomain pointing at the old server too. Left alone, it would have been a candidate for takeover once the old address was released back into the provider's pool, so it went the same way as the AAAA. The SPF builder is a handy way to rewrite the TXT record once you know who really sends your mail.

What the old server was doing

It is worth asking why the old address did anything at all. The previous plan had been cancelled in principle but not yet terminated, so the machine stayed up for the rest of the month. Its web server, no longer holding the site's files, answered with its default configuration: a generic placeholder page for some requests and a certificate error for others, since the site's certificate had been removed along with the account.

Visitors who got the certificate warning mostly gave up, and those who got the placeholder assumed the charity had closed. That is a reputational cost in its own right, well beyond the lost page views. After the contract ended, the address would have started timing out instead, which would have changed the symptom from "wrong page" to "slow page" without fixing anything. A fault that changes its character as the old server decays makes it harder still to describe in a ticket.

Verifying it

After any move, list every record in the zone and ask of each one: is this still true? Look specifically for AAAA, MX, TXT and any service subdomains that pointed at the old host.

What would have caught it

PreviousThe HSTS setting that broke the intranetNextThe mailbox that filled up and bounced the best client

More from Hosting Autopsy

Autopsy

The HSTS setting that broke the intranet

After a security review, an administrator enabled HSTS on the company's main domain with the option that...

Autopsy

The autoscaler that scaled to meet the bots

A startup used auto-scaling so the site would handle spikes. During one weekend a scraper began requesting...

Autopsy

The plugin that emailed every order to someone who had left

Years earlier, a shop's order notification plugin had been configured to send a copy of every order to the...