An organisation used one wildcard certificate for its website and also installed the same files on its mail server. The web side renewed itself automatically. The mail server was a separate machine where someone had copied the files by hand two years earlier.
On the day the old certificate expired, mail clients across the office began showing warnings and some refused to connect. Phones that synchronised mail stopped syncing. Nobody connected it to the website certificate for two hours, because the website looked fine.
They set up automatic deployment of the renewed certificate to every service that used it and made a list of where each one lives. The organisation in this account, a regional charity with about forty staff, is a composite. The failure is an old one: a certificate is a file, and a file copied by hand does not renew itself.
One certificate, several homes
A wildcard certificate covers every name at one level beneath a domain. One issued for *.example.com is valid for www.example.com, mail.example.com, shop.example.com and any other single label in that position. It does not cover the bare example.com unless that name is listed too, and it does not cover deeper names such as a.b.example.com.
That coverage is what tempted the charity's first administrator. One purchase and one set of files would serve everything. He installed the certificate on the web server, then copied the same certificate file and private key to the mail server so that staff could connect to mail.example.com over an encrypted channel.
Those copies are the problem. The certificate does not belong to one server. It exists wherever somebody put a copy, and each copy has the same expiry date and nothing to refresh it.
Why the website renewed and the mail server did not
Since the charity moved its site to a modern setup, renewal on the web server had been automatic. An ACME client, the same kind of program used by most certificate authorities these days, ran from a scheduled task, contacted the authority about a month before expiry, proved control of the domain and installed the new files. Because it was a wildcard, the proof was done through a DNS record, which the client created through an API.
The client knew about the web server and reloaded it after each renewal. It had never been told about the mail server, which was a different machine in a different place, run by a different person at the time. There was nothing connecting the two. The mail server carried on using a file that had been copied once, two years earlier, and which nobody thought of as a thing that expires.
That is why the website looked perfect on the day. The padlock was fine. The renewal had worked, just as designed. The design had a hole in it that was not visible from any browser.
The morning it broke
At around 08:30 on a Monday, staff opened their mail programs. Several saw a warning that the server's certificate had expired, with the choice of continuing or cancelling. The cautious ones cancelled and rang the helpdesk. Others clicked through, which is exactly the habit a certificate warning is meant to discourage.
Some clients were stricter. Phones that synchronised mail in the background gave up without showing anything, so staff discovered the problem only when they noticed that no new messages had arrived since the weekend. A few automated systems that sent mail through the server, such as the donation receipt system, began to log errors about failed TLS handshakes.
The helpdesk suspected the mail server's software and spent the first hour restarting services. Then they suspected a password policy change. The certificate was a late suspect because it had been renewed, they thought, two weeks earlier, and the website padlock confirmed it. Checking the right server finally took two hours to occur to anyone.
What a wildcard does and does not cover
The charity's staff assumed the certificate "covered the domain", which is close but not exact. The table shows how a certificate for *.example.com treats several names, and it explains why the administrator had also listed the bare name when he ordered it.
| Name | Covered by *.example.com? | Why |
|---|---|---|
| www.example.com | Yes | One label in the wildcard position |
| mail.example.com | Yes | One label in the wildcard position |
| example.com | Only if listed separately | No label to match the asterisk |
| a.b.example.com | No | Two labels, the wildcard matches one |
None of this was the cause of the outage, but it shaped the response. When the replacement was installed, the team checked that the mail client's configured server name, mail.example.com, matched the certificate. A staff phone configured with an older name, imap.example.com, would have matched too, but one set to the server's raw hostname would not, and two such phones kept complaining after everyone else was fixed.
How it was diagnosed
Once someone looked at the mail server's own certificate, it took seconds. The mail ports each have their own handshake, and openssl can speak to all of them:
# IMAP over TLS
echo | openssl s_client -connect mail.example.com:993 -servername mail.example.com 2>/dev/null | openssl x509 -noout -subject -enddate
# SMTP submission with STARTTLS
echo | openssl s_client -starttls smtp -connect mail.example.com:587 2>/dev/null | openssl x509 -noout -enddate
Both printed a notAfter date from that morning. A comparison against the web server's date, a year ahead, showed two certificates with the same subject and different lifetimes. The mail services were presenting the old one.
The fix, and the better fix
The quick fix was to copy the renewed files to the mail server and reload both services. In a typical Postfix and Dovecot setup the reload is two commands:
systemctl reload postfix
systemctl reload dovecot
Reloading matters. Both programs read the certificate when they start, so replacing the file on disk achieves nothing until they are told to read it again.
The better fix removed the manual step. The ACME client was given a deploy hook, a script that runs after each successful renewal. This one copies the new certificate and key to the mail server over a restricted SSH account, sets the permissions and reloads the services:
#!/bin/sh
# /etc/letsencrypt/renewal-hooks/deploy/push-mail.sh
scp /etc/ssl/example.com/fullchain.pem [email protected]:/etc/ssl/example.com/
scp /etc/ssl/example.com/privkey.pem [email protected]:/etc/ssl/example.com/
ssh [email protected] 'systemctl reload postfix dovecot'
The alternative was to give the mail server its own certificate for mail.example.com, issued and renewed by its own client. That is often the cleaner design, as it keeps the private key of a wildcard off every machine that does not need it. A wildcard key on a mail server means a compromise of the mail server lets someone impersonate every name in the domain.
Things people ask
Does an expired certificate stop mail between servers?
Often not. Server-to-server SMTP on port 25 is usually opportunistic, and many senders carry on regardless. Strict policies, such as MTA-STS, change that, and so does your own staff's mail client.
Is a wildcard certificate a bad idea?
Not in itself. It simplifies names, but the key then needs to be held carefully and distributed deliberately.
Can the same certificate cover web and mail?
Yes, if the names match, but treat each place it is installed as an item that needs renewal.
How early should monitoring warn?
Thirty days, then fourteen, then seven, so that a missed renewal has time to be noticed.
What would have caught it
- Inventory every place a certificate is installed, mail servers included.
- Automate distribution to secondary services or give each its own certificate.
- Monitor mail, FTP and other ports for expiry too, since 443 is only one of them.
- Reload the services after replacing the files, and test by connecting, not just by checking that the file is on disk.