Hosting Autopsy / The certificate on one server, but not its twin

The certificate on one server, but not its twin

HOSTING AUTOPSY

8 min read · 1,678 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A company ran its site behind a load balancer with two identical web servers. When the certificate was renewed, the administrator installed it on the first server and the load balancer, and forgot the second. It is the sort of omission that takes thirty seconds to commit and several days to find, because nothing about it looks broken from the front door.

Half of the requests were routed to the second server, which still presented the expired certificate for the internal connection. Because the load balancer was lenient about verification, the site mostly worked, but certain tools and API clients began failing at random.

The pattern, "it works every other time", made it hard to track down, until someone tested each server directly. This is a composite story, built from several real incidents, but the shape of it will be familiar to anyone who has run more than one machine behind a single address.

The setup: one address, two twins

The company was a mid-sized business selling to trade customers, with a public site and a small API that a few partners called from their own systems. In front sat a load balancer holding the public address. Behind it were two web servers, built from the same image, running the same application and the same web server software. The idea was that if one machine needed patching, the other carried on.

Traffic reached the servers over HTTPS even on the private network between them. The load balancer accepted the visitor's connection, decrypted it, picked a server, and opened a fresh encrypted connection to that server. This is usually called re-encryption, and it is common where policy says traffic should not cross even an internal network in the clear.

That means three certificates were in play, not one: the one the load balancer showed to visitors, and one on each web server. Visitors only ever saw the first. The other two lived on the inner leg of the journey, where nobody was looking.

VisitorLoad balancercert: validServer 1cert: renewedServer 2cert: expiredHTTPSinner HTTPS, not verified
The visitor only ever sees the balancer's certificate; the two inner connections carry their own.

What the renewal missed

The renewal itself was routine. The old certificate was approaching its end date, a new one was issued, and the administrator worked through the installation by hand: upload the files to the first server, reload the web server, then upload the new certificate to the load balancer's front listener. The check at the end was to open the site in a browser. The padlock was fine, the page loaded, and the ticket was closed.

The second server was not in the administrator's head. It had been built months earlier by a colleague who had since moved teams, and the runbook for renewals listed "web server" in the singular. Nothing in the process asked how many web servers there were.

It is worth being clear about why the browser check proved nothing. A browser tests the certificate on the outer leg, the balancer's. The inner leg is invisible to it. The check confirmed exactly one of the three certificates and said nothing about the other two.

Why it mostly worked

When the old certificate passed its end date, the load balancer carried on regardless. It had been configured, as many balancers are, to encrypt the inner connection without verifying what came back. No check of the expiry date, no check of the name, no check of the issuer. The encryption was still real; the identity check was switched off. Whoever set it up had done so years earlier to get past a self-signed certificate and never turned it back on.

So for ordinary visitors, nothing happened. Their browser talked to the balancer, the balancer talked to either server, and the expired certificate on the second machine went unchallenged. Page views, basket, checkout: all fine.

The trouble came from the clients that did not go through the front door in the normal way. The partners' API traffic reached the servers through a second listener on the balancer set to pass TCP straight through without decrypting it. Those clients saw whichever server's certificate the balancer picked, and strict API libraries refuse an expired one. So did a monitoring probe and a reporting tool that connected to the server pool by its internal name.

Round-robin balancing sent connections alternately to each server. One request in two got the good certificate, and the next got the bad one.

OKfailOKfailOKfailOKfailsrv 1srv 2srv 1srv 2srv 1srv 2srv 1srv 2Passthrough API requests, in order (illustrative)
With round-robin balancing, a request in two reached the server with the expired certificate.

The wrong turns

The first reports came from a partner, who said the API "keeps timing out" and, a day later, that it "works if you retry". That phrasing sent the team straight to the application. They read logs, found nothing odd, and restarted the application servers. The failures continued, and the restart made things look better for an hour because a quiet hour has fewer requests to fail.

Next they suspected the partner's network, then the firewall, then a recent library update on the partner's side. Each theory had a grain of plausibility and none explained why the same client, from the same place, succeeded and failed in alternation. The pattern was in the data from the start. Failures at precisely every second request are not weather; they are arithmetic.

The most expensive wrong turn was blaming the balancer's health checks. They were passing, and they would keep passing, because health checks on this setup used plain HTTP to a status page and never touched the certificate at all.

Finding it: asking each server directly

The breakthrough came from a support engineer who ignored the balancer entirely and connected to each server by its own address. The command is short:

echo | openssl s_client -connect 203.0.113.11:443 -servername www.example.com 2>/dev/null | openssl x509 -noout -enddate
echo | openssl s_client -connect 203.0.113.12:443 -servername www.example.com 2>/dev/null | openssl x509 -noout -enddate

The first printed a date a few months ahead. The second printed a date nine days in the past. That was the whole case, in two lines of output.

A quick confirmation with curl made the failure visible in the client's own words:

curl -sv --resolve www.example.com:443:203.0.113.12 https://www.example.com/ -o /dev/null
* SSL certificate problem: certificate has expired

Using --resolve pins the hostname to a chosen address without touching DNS, which makes it a handy way to test one machine at a time. The point was to take the balancer out of the conversation. A balancer is designed to smooth over differences between its back ends; a diagnosis needs those differences exposed.

The fix and the aftermath

The repair took ten minutes: copy the renewed certificate and key to the second server, reload its web server, and re-run the two commands above. Both now showed the same end date. The partner's retries stopped being necessary the same afternoon.

The tidying-up took longer. The team turned inner-leg verification back on at the balancer, after first confirming that every server presented a certificate with the right name and a trusted issuer. That change would have made the original mistake loud within minutes instead of silent for nine days.

They also replaced the hand-copied process. Certificates are now issued by an automated client on a single machine and pushed to every server and to the balancer by a script that reads the server list from one file, so adding a third server means editing one line. The runbook says "every server", and the script logs each one it touches. Finally, the expiry monitor was extended to cover each back end by address, not only the public name.

Turning verification on after years of leniency can itself cause an outage if a back end has a self-signed or mismatched certificate. Check every server first, with the commands above, then change the setting.

Checking it yourself

If you run anything behind a balancer, list the machines that terminate TLS: the balancer itself, each back end that speaks HTTPS, and any passthrough listener. For each one, run the openssl s_client command above against its own address and read the expiry date and the subject.

  1. Find every address behind the public name, from the balancer's pool configuration, not from memory.
  2. Query each address directly with the right -servername, and compare end dates and serial numbers.
  3. Test the public name several times in a row; a pattern of alternate failures points at one bad back end.
  4. Confirm whether the balancer verifies the inner connection, and what it does when verification fails.

The SSL expiry tool helps with the public name, and the troubleshooting guide has the wider checklist.

Loose ends

Should the inner connection be encrypted at all?

It depends on your policy and who shares the network. On a single-tenant private network many sites terminate TLS at the balancer and use plain HTTP inside. If you do re-encrypt, verify properly; unverified encryption protects against passive listeners but not impostors.

Can one certificate be shared by all servers?

Yes, and it is common: the same certificate and key on every server for the same hostname. It makes renewal a single issue followed by distribution, provided the distribution is automated.

Why did the health checks stay green?

Because they tested the application over HTTP, not the certificate. A health check only catches what it exercises.

Would a short-lived certificate have helped?

Short lifetimes make automation mandatory, which is the real cure, but they also make a forgotten server fail sooner.

What would have caught it

PreviousThe mailbox that filled up and bounced the best clientNextThe domain contact who no longer worked there

More from Hosting Autopsy

Autopsy

The sitemap that listed forty thousand junk URLs

A furniture retailer added a filter system to its catalogue: colour, material, price band, size. Each...

Autopsy

The CAA record that blocked its own certificate

A security-minded administrator added a CAA record listing the one authority the company had used for years....

Autopsy

The launch-day database that ran out of connections

A local theatre put tickets for its autumn season on sale at noon and announced it on social media a day...