A hobbyist forum, the kind run by one volunteer admin for a community of a few thousand people interested in the same obscure subject, had a nightly backup running faithfully for over a year. Every morning the control panel showed a green tick and a file with a recent date. The admin glanced at it now and then, felt reassured, and got on with moderating.
Then a badly written plugin corrupted the database. Topics vanished, then the tables behind them. The admin did what everyone is told to do: went to restore from last night's backup.
The archive opened fine. The files were all there. The database dump, however, was zero bytes. This is a composite story, but every detail in it is a pattern I have met in real tickets, and the last one still makes me wince.
What had actually been running
The backup was a small script run from cron at 03:00. It did two jobs in sequence: dump the database into a file with mysqldump, then pack the site files and that dump into a single compressed archive and store it in a backups folder. A second job pruned archives older than fourteen days.
#!/bin/sh
cd /home/forum
mysqldump -u forum_user -p"$DBPASS" forum_db > backups/db.sql
tar czf backups/forum-$(date +%F).tar.gz public_html backups/db.sql
find backups -name 'forum-*.tar.gz' -mtime +14 -delete
Read it again and look for the line that checks whether anything worked. There is not one. If mysqldump fails, the shell carries on to the next line, because nothing told it not to. The tar command then happily packs an empty file next to several gigabytes of site files and exits with success. Cron sees exit code zero and is satisfied. The control panel sees a new file and draws a tick.
The tidy-up that broke it
Some months earlier, the admin had done a routine clean-up of the database server: removing old test users, tightening privileges, trimming accounts that looked unused. The backup used a database user whose permissions had been narrowed along the way. The account could still read and write the forum's tables for normal operation but had lost the extra privileges mysqldump needs, such as the ability to lock tables and to read views and routines.
From that night, every dump failed with an error. The error went to standard error, which cron emails to the account owner, and the account owner's address was an old mailbox that nobody read. So the failure message was produced every single night for months and never seen by anyone.
What the file sizes would have said
The signal was there in plain sight. A forum of that age would have had a dump of around 2 GB, compressing to well under 1 GB. After the permissions change, the archive became files only, a small fraction of that size. Nobody compared them.
Restore day
When the plugin corruption hit, the admin took the forum offline and unpacked last night's archive. Files fine. Then ls -l backups/db.sql showed zero bytes. The next night's was also zero. So was the one before that. The admin worked back through two weeks of retained archives, all empty in the same way, and the fourteen-day pruning job had already deleted anything older.
There was no second copy. The hosting plan included weekly account snapshots, but those lived in the same account, and the one tried first turned out to be a snapshot of files only. The control panel's restore button offered no database option at all, which the admin had never noticed. The corrupted database itself could not be repaired; the damaged tables had been overwritten.
What was recovered
Very little. Eight years of conversations were gone. A partial copy of some of the more popular threads existed in a search engine's cache and in the Internet's various web archives, and a few members had saved pages of their own posts. The admin spent two weeks pasting those back by hand and reloaded the member list from an old mailing export.
The forum limped on, but most of its long-term members never returned. For a community site, the history is the product. Take away the archive and what remains is a new, empty forum with an old name.
The fix, and what to do instead
The technical fix is small and the discipline is where the work lies. Turn on failure on error in the script so it stops at the first problem, check the size of the dump before packing, and send the result somewhere a person will read it.
#!/bin/sh
set -e
cd /home/forum
mysqldump --single-transaction -u forum_user -p"$DBPASS" forum_db > backups/db.sql
[ "$(stat -c %s backups/db.sql)" -gt 1000000 ] || { echo "dump too small" >&2; exit 1; }
tar czf backups/forum-$(date +%F).tar.gz public_html backups/db.sql
rclone copy backups/forum-$(date +%F).tar.gz offsite:forum-backups/
Notice the last line: a copy leaves the hosting account altogether. If the account is compromised, suspended or deleted, local backups go with it. Then give the cron job a working notification address and use a database user with exactly the privileges a dump requires, documented in a comment.
Other ways a backup lies
This particular failure, an empty dump inside a healthy-looking archive, is one of a family. The common thread is that something reports on its own step and nobody reports on the whole.
| Failure | What you see | How to spot it |
|---|---|---|
| Empty or truncated dump | Green tick, small archive | Size check, tail of the dump for the closing comment line |
| Disk full during the job | Archive exists but is cut short | gzip -t or tar tzf reports an unexpected end |
| Changed password | Dump fails every night | Mail from cron, if anyone reads it |
| Wrong database name after a migration | Backs up an old, empty database | Restore test |
| Backups stored on the same disk | Everything fine until the disk dies | Ask where the copies physically live |
A correctly finished mysqldump file ends with a line starting "-- Dump completed on". Checking for that line takes one command and catches the truncated case as well as the empty one. For anything bigger than a hobby forum, the same idea applies to every export: verify the result, not the attempt.
Checking it yourself
Once a quarter, restore. Create an empty scratch database, load the latest dump, and open the restored site on a test address. Look at the newest post and the oldest one, log in as a real member, and try a search. Write down how long the whole thing took, because that figure is your real recovery time, and it is the number to quote when someone asks how quickly you could be back. If you cannot do that in an hour, your recovery plan is slower than you think.
mysql -u scratch -p scratch_db < db.sql
mysql -u scratch -p -e "SELECT COUNT(*), MAX(created) FROM scratch_db.posts"
Between restores, check the backup size every day against the previous week and alert on a drop of more than a fifth. Run tar tzf on the archive and make sure the dump is listed with a plausible size. A tick on a dashboard says that something happened; it does not say that the thing was any use. The troubleshooting guide covers a few more of these habits.
Loose ends
Is the host's backup enough?
It is a useful second copy, not a replacement. Know how often it runs, how long it is kept and whether it is stored away from your account.
How many copies should I keep?
The old rule of three copies on two kinds of media with one off-site still holds. In practice: the live site, a nightly archive on the server, and a daily copy elsewhere.
Does a successful restore test prove everything?
It proves the last one worked. Repeat it after any change to users, servers or the backup script.
What would have caught it
- Restoring a backup to a scratch location every quarter. It takes an hour and it is the only real test.
- Alerts on backup size: a sudden drop is a signal, and so is a total that never changes.
- Keeping at least one copy outside the hosting account, so a compromise or suspension does not take the backups too.
- A script that stops on the first error and sends mail to an address that is actually read.
- A note in the privilege clean-up checklist saying which accounts the backup depends on.