Hosting Autopsy / The backup that could not be restored

The backup that could not be restored

HOSTING AUTOPSY

7 min read · 1,439 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A hobbyist forum, the kind run by one volunteer admin for a community of a few thousand people interested in the same obscure subject, had a nightly backup running faithfully for over a year. Every morning the control panel showed a green tick and a file with a recent date. The admin glanced at it now and then, felt reassured, and got on with moderating.

Then a badly written plugin corrupted the database. Topics vanished, then the tables behind them. The admin did what everyone is told to do: went to restore from last night's backup.

The archive opened fine. The files were all there. The database dump, however, was zero bytes. This is a composite story, but every detail in it is a pattern I have met in real tickets, and the last one still makes me wince.

What had actually been running

The backup was a small script run from cron at 03:00. It did two jobs in sequence: dump the database into a file with mysqldump, then pack the site files and that dump into a single compressed archive and store it in a backups folder. A second job pruned archives older than fourteen days.

#!/bin/sh
cd /home/forum
mysqldump -u forum_user -p"$DBPASS" forum_db > backups/db.sql
tar czf backups/forum-$(date +%F).tar.gz public_html backups/db.sql
find backups -name 'forum-*.tar.gz' -mtime +14 -delete

Read it again and look for the line that checks whether anything worked. There is not one. If mysqldump fails, the shell carries on to the next line, because nothing told it not to. The tar command then happily packs an empty file next to several gigabytes of site files and exits with success. Cron sees exit code zero and is satisfied. The control panel sees a new file and draws a tick.

Dump database (mysqldump) Permission denied, 0 bytes Pack files + empty dump Panel shows green tick Nothing in the chain checks the size of the dump or the exit code of the step before it.
The failure happens at step one, but every later step only checks that it managed to do its own job.

The tidy-up that broke it

Some months earlier, the admin had done a routine clean-up of the database server: removing old test users, tightening privileges, trimming accounts that looked unused. The backup used a database user whose permissions had been narrowed along the way. The account could still read and write the forum's tables for normal operation but had lost the extra privileges mysqldump needs, such as the ability to lock tables and to read views and routines.

From that night, every dump failed with an error. The error went to standard error, which cron emails to the account owner, and the account owner's address was an old mailbox that nobody read. So the failure message was produced every single night for months and never seen by anyone.

What the file sizes would have said

The signal was there in plain sight. A forum of that age would have had a dump of around 2 GB, compressing to well under 1 GB. After the permissions change, the archive became files only, a small fraction of that size. Nobody compared them.

Archive size by month, illustrative (arbitrary units) DB dump missing from here on before after the permission change
The archive shrank by more than half overnight and stayed there for months, with the tick still green.

Restore day

When the plugin corruption hit, the admin took the forum offline and unpacked last night's archive. Files fine. Then ls -l backups/db.sql showed zero bytes. The next night's was also zero. So was the one before that. The admin worked back through two weeks of retained archives, all empty in the same way, and the fourteen-day pruning job had already deleted anything older.

There was no second copy. The hosting plan included weekly account snapshots, but those lived in the same account, and the one tried first turned out to be a snapshot of files only. The control panel's restore button offered no database option at all, which the admin had never noticed. The corrupted database itself could not be repaired; the damaged tables had been overwritten.

What was recovered

Very little. Eight years of conversations were gone. A partial copy of some of the more popular threads existed in a search engine's cache and in the Internet's various web archives, and a few members had saved pages of their own posts. The admin spent two weeks pasting those back by hand and reloaded the member list from an old mailing export.

The forum limped on, but most of its long-term members never returned. For a community site, the history is the product. Take away the archive and what remains is a new, empty forum with an old name.

The fix, and what to do instead

The technical fix is small and the discipline is where the work lies. Turn on failure on error in the script so it stops at the first problem, check the size of the dump before packing, and send the result somewhere a person will read it.

#!/bin/sh
set -e
cd /home/forum
mysqldump --single-transaction -u forum_user -p"$DBPASS" forum_db > backups/db.sql
[ "$(stat -c %s backups/db.sql)" -gt 1000000 ] || { echo "dump too small" >&2; exit 1; }
tar czf backups/forum-$(date +%F).tar.gz public_html backups/db.sql
rclone copy backups/forum-$(date +%F).tar.gz offsite:forum-backups/

Notice the last line: a copy leaves the hosting account altogether. If the account is compromised, suspended or deleted, local backups go with it. Then give the cron job a working notification address and use a database user with exactly the privileges a dump requires, documented in a comment.

Other ways a backup lies

This particular failure, an empty dump inside a healthy-looking archive, is one of a family. The common thread is that something reports on its own step and nobody reports on the whole.

FailureWhat you seeHow to spot it
Empty or truncated dumpGreen tick, small archiveSize check, tail of the dump for the closing comment line
Disk full during the jobArchive exists but is cut shortgzip -t or tar tzf reports an unexpected end
Changed passwordDump fails every nightMail from cron, if anyone reads it
Wrong database name after a migrationBacks up an old, empty databaseRestore test
Backups stored on the same diskEverything fine until the disk diesAsk where the copies physically live

A correctly finished mysqldump file ends with a line starting "-- Dump completed on". Checking for that line takes one command and catches the truncated case as well as the empty one. For anything bigger than a hobby forum, the same idea applies to every export: verify the result, not the attempt.

Checking it yourself

Once a quarter, restore. Create an empty scratch database, load the latest dump, and open the restored site on a test address. Look at the newest post and the oldest one, log in as a real member, and try a search. Write down how long the whole thing took, because that figure is your real recovery time, and it is the number to quote when someone asks how quickly you could be back. If you cannot do that in an hour, your recovery plan is slower than you think.

mysql -u scratch -p scratch_db < db.sql
mysql -u scratch -p -e "SELECT COUNT(*), MAX(created) FROM scratch_db.posts"

Between restores, check the backup size every day against the previous week and alert on a drop of more than a fifth. Run tar tzf on the archive and make sure the dump is listed with a plausible size. A tick on a dashboard says that something happened; it does not say that the thing was any use. The troubleshooting guide covers a few more of these habits.

Loose ends

Is the host's backup enough?

It is a useful second copy, not a replacement. Know how often it runs, how long it is kept and whether it is stored away from your account.

How many copies should I keep?

The old rule of three copies on two kinds of media with one off-site still holds. In practice: the live site, a nightly archive on the server, and a daily copy elsewhere.

Does a successful restore test prove everything?

It proves the last one worked. Repeat it after any change to users, servers or the backup script.

What would have caught it

PreviousThe plugin update that took down the storeNextThe contact form that blacklisted the server

More from Hosting Autopsy

Autopsy

The twenty-four hour TTL on the day the server died

A site's DNS records had a TTL of 86,400 seconds, a full day, set years ago by a default. One morning the...

Autopsy

The contact form that mailed an expired domain

A small manufacturer's contact form sent enquiries to an address at a domain the owner had used years ago for...

Autopsy

The certificate that expired on a Friday evening

An online shop selling hand-made furniture ran for three years on a certificate bought once and renewed by...