A VPS hosted four or five small sites without problems for two years. One morning all of them returned errors, and the database would not start. Logging in over SSH, the administrator saw the cause in one command: the disk was at 100 percent.
The culprit was a log file. A misbehaving script had been writing an error message hundreds of times a second after a software update, and the log rotation, which should have compressed and trimmed old logs, had been set up for a different path. By the time anyone noticed, one file had grown to over 40 gigabytes.
A full disk is unkind to databases, which cannot write their transaction logs and may stop or even corrupt data if they crash at the wrong moment. Here, luckily, deleting the log and restarting the services was enough.
What the morning looked like
The server belonged to a freelance developer who hosted her own portfolio and the sites of a few clients: a bakery, a walking club, a two-person design studio and a small charity. It had 80 GB of disk, a modest amount of RAM and a calm history. Updates were applied monthly, backups ran nightly, and the last real incident had been a certificate renewal in the first year.
At about seven in the morning the bakery owner messaged to say the order form was giving an error. By half past, the walking club's site was down too, and the developer's own site showed a blank white page. Several sites failing together and in different ways pointed at something shared, so she skipped the websites and went straight to the server.
$ df -h /
Filesystem Size Used Avail Use% Mounted on
/dev/vda1 79G 79G 0 100% /
Zero available. That one line explained every symptom, and it is the first command to run whenever several unrelated sites fail at once.
Why a full disk breaks everything at once
Nearly every program on a server writes something as it runs. Web servers write access logs, PHP writes sessions and temporary upload files, mail systems write queues, and databases write their data files and a write-ahead log or redo log that records changes before they are applied. When there is no room, those writes fail, and the failures show up as whatever each program does when it cannot write: a blank page, a login that does not stick, an upload that vanishes, or a service that refuses to start.
Databases are the sharp edge. MariaDB and MySQL need space for temporary tables and for transaction logs. When the disk fills mid-write, the database usually stops itself cleanly and complains in its error log. If the machine crashes or the process is killed at the wrong instant, however, the tables can be left half-written and need a repair. That is the corruption risk the administrator was thinking of, and why the first priority was freeing space before restarting anything.
Finding the culprit
With the disk full, the question is what is big. The tool for that is du, working down from the root one level at a time and keeping to one filesystem.
du -xh --max-depth=1 / 2>/dev/null | sort -h | tail
du -xh --max-depth=1 /var 2>/dev/null | sort -h | tail
ls -lhS /var/log | head
The first command pointed at /var, the second at /var/log, and the third gave a name: a PHP error log for one of the client sites, 43 GB. A look at its tail showed the same warning repeated line after line, every line a few milliseconds after the last.
The warning came from a plugin that had been updated the previous evening. It called a function that no longer existed in the PHP version running on the server, hit the failure on every page request, and logged it every time. With a modest number of visitors and a few thousand requests from search crawlers, it reached tens of gigabytes in about thirteen hours.
Freeing space the right way
Here is where a common mistake is easy to make. Deleting a log file that a running program still has open does not free the space. The file's name disappears, but the data stays on disk until the program closes it, so df keeps reporting a full disk and the administrator concludes that nothing worked. The reliable approach is to empty the file in place:
truncate -s 0 /var/log/php/clientsite-error.log
df -h /
This releases the space at once and the program carries on writing to the same file. If you have already deleted something and space has not come back, lsof +L1 lists files that are deleted but still held open, and restarting the process that owns them lets go.
The developer truncated the file, then restarted MariaDB, checked its error log for crash recovery messages, and ran a check on the tables the bakery's site used. Everything was intact. Only then did she turn to the cause.
The rotation that was not rotating
Logrotate's rules live in /etc/logrotate.d/, one file per application, each naming the paths it covers. When the developer moved each site's logs into per-site folders the year before, she updated the web server's configuration but not the matching logrotate file. Nothing complained, because a rule that matches no files is not an error.
The fix was to add the new path to the rule, then test it without waiting for the nightly run. The dry-run flag prints what logrotate would do:
logrotate -d /etc/logrotate.d/php-sites
logrotate -f /etc/logrotate.d/php-sites
The second command forces a rotation so you can see a compressed .gz file appear. Then she patched the plugin by rolling it back, which stopped the flood at the source. Clearing the log without fixing the cause would have refilled the disk in half a day.
The aftermath
The outage lasted about two hours from first complaint to everything answering again. Orders placed on the bakery site in that window were lost, because the checkout could not write them, and the owner re-entered some from phone calls. The developer wrote a short note to all four clients describing what had failed, what was fixed and what she was changing, without jargon and without blaming the plugin vendor.
She then added three things. A disk alert at 80 percent that emails her and a backup contact. A weekly cron job that lists the ten largest files under /var. And a line in her update routine: after any software update, tail the error logs for a minute to see whether something has started complaining at speed.
What changes by hosting type
| Type | What fills up | Who watches it |
|---|---|---|
| Shared hosting | Account quota, mostly mail and error logs | The host warns; you clean up |
| VPS | The whole disk, including the OS | You, or nobody |
| Dedicated server | Disks, sometimes in a RAID set | You, plus hardware alerts |
| Managed hosting | Plan storage allowance | The provider, usually with alerts |
The VPS row is the dangerous one. You get the freedom of a whole server and none of the guard rails of a managed plan, and a machine that has behaved for two years teaches you to stop looking.
What would have caught it
- A monitor that warns when disk usage passes 80 percent.
- Log rotation tested on every log path actually in use.
- Fixing the error in the script as well as clearing the log. It would otherwise fill the disk again.
- A weekly look at the five biggest directories, so growth shows up as a trend before it becomes an outage.