Hosting Autopsy / The disk that filled up overnight

The disk that filled up overnight

HOSTING AUTOPSY

6 min read · 1,291 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A VPS hosted four or five small sites without problems for two years. One morning all of them returned errors, and the database would not start. Logging in over SSH, the administrator saw the cause in one command: the disk was at 100 percent.

The culprit was a log file. A misbehaving script had been writing an error message hundreds of times a second after a software update, and the log rotation, which should have compressed and trimmed old logs, had been set up for a different path. By the time anyone noticed, one file had grown to over 40 gigabytes.

A full disk is unkind to databases, which cannot write their transaction logs and may stop or even corrupt data if they crash at the wrong moment. Here, luckily, deleting the log and restarting the services was enough.

What the morning looked like

The server belonged to a freelance developer who hosted her own portfolio and the sites of a few clients: a bakery, a walking club, a two-person design studio and a small charity. It had 80 GB of disk, a modest amount of RAM and a calm history. Updates were applied monthly, backups ran nightly, and the last real incident had been a certificate renewal in the first year.

At about seven in the morning the bakery owner messaged to say the order form was giving an error. By half past, the walking club's site was down too, and the developer's own site showed a blank white page. Several sites failing together and in different ways pointed at something shared, so she skipped the websites and went straight to the server.

$ df -h /
Filesystem      Size  Used Avail Use% Mounted on
/dev/vda1        79G   79G     0 100% /

Zero available. That one line explained every symptom, and it is the first command to run whenever several unrelated sites fail at once.

Why a full disk breaks everything at once

Nearly every program on a server writes something as it runs. Web servers write access logs, PHP writes sessions and temporary upload files, mail systems write queues, and databases write their data files and a write-ahead log or redo log that records changes before they are applied. When there is no room, those writes fail, and the failures show up as whatever each program does when it cannot write: a blank page, a login that does not stick, an upload that vanishes, or a service that refuses to start.

Databases are the sharp edge. MariaDB and MySQL need space for temporary tables and for transaction logs. When the disk fills mid-write, the database usually stops itself cleanly and complains in its error log. If the machine crashes or the process is killed at the wrong instant, however, the tables can be left half-written and need a repair. That is the corruption risk the administrator was thinking of, and why the first priority was freeing space before restarting anything.

Disk use, 80 GB total (illustrative)NormalSites + DBSystemIncidentSites + DBSystemOne log file, 40 GB+Everything else stayed the same size; one file took all the free space.
Normal use sat below half the disk; a single runaway file consumed the rest.

Finding the culprit

With the disk full, the question is what is big. The tool for that is du, working down from the root one level at a time and keeping to one filesystem.

du -xh --max-depth=1 / 2>/dev/null | sort -h | tail
du -xh --max-depth=1 /var 2>/dev/null | sort -h | tail
ls -lhS /var/log | head

The first command pointed at /var, the second at /var/log, and the third gave a name: a PHP error log for one of the client sites, 43 GB. A look at its tail showed the same warning repeated line after line, every line a few milliseconds after the last.

The warning came from a plugin that had been updated the previous evening. It called a function that no longer existed in the PHP version running on the server, hit the failure on every page request, and logged it every time. With a modest number of visitors and a few thousand requests from search crawlers, it reached tens of gigabytes in about thirteen hours.

Freeing space the right way

Here is where a common mistake is easy to make. Deleting a log file that a running program still has open does not free the space. The file's name disappears, but the data stays on disk until the program closes it, so df keeps reporting a full disk and the administrator concludes that nothing worked. The reliable approach is to empty the file in place:

truncate -s 0 /var/log/php/clientsite-error.log
df -h /

This releases the space at once and the program carries on writing to the same file. If you have already deleted something and space has not come back, lsof +L1 lists files that are deleted but still held open, and restarting the process that owns them lets go.

The developer truncated the file, then restarted MariaDB, checked its error log for crash recovery messages, and ran a check on the tables the bakery's site used. Everything was intact. Only then did she turn to the cause.

Rotation rule watches/var/log/php/*.logRule for old setup,files no longer thereActual log being written/var/log/sites/clientsite.lognever matches
A rotation rule that points at the wrong path looks fine in the configuration and does nothing.

The rotation that was not rotating

Logrotate's rules live in /etc/logrotate.d/, one file per application, each naming the paths it covers. When the developer moved each site's logs into per-site folders the year before, she updated the web server's configuration but not the matching logrotate file. Nothing complained, because a rule that matches no files is not an error.

The fix was to add the new path to the rule, then test it without waiting for the nightly run. The dry-run flag prints what logrotate would do:

logrotate -d /etc/logrotate.d/php-sites
logrotate -f /etc/logrotate.d/php-sites

The second command forces a rotation so you can see a compressed .gz file appear. Then she patched the plugin by rolling it back, which stopped the flood at the source. Clearing the log without fixing the cause would have refilled the disk in half a day.

The aftermath

The outage lasted about two hours from first complaint to everything answering again. Orders placed on the bakery site in that window were lost, because the checkout could not write them, and the owner re-entered some from phone calls. The developer wrote a short note to all four clients describing what had failed, what was fixed and what she was changing, without jargon and without blaming the plugin vendor.

She then added three things. A disk alert at 80 percent that emails her and a backup contact. A weekly cron job that lists the ten largest files under /var. And a line in her update routine: after any software update, tail the error logs for a minute to see whether something has started complaining at speed.

What changes by hosting type

TypeWhat fills upWho watches it
Shared hostingAccount quota, mostly mail and error logsThe host warns; you clean up
VPSThe whole disk, including the OSYou, or nobody
Dedicated serverDisks, sometimes in a RAID setYou, plus hardware alerts
Managed hostingPlan storage allowanceThe provider, usually with alerts

The VPS row is the dangerous one. You get the freedom of a whole server and none of the guard rails of a managed plan, and a machine that has behaved for two years teaches you to stop looking.

What would have caught it

PreviousThe redirect loop nobody could seeNextThe sitemap that listed forty thousand junk URLs

More from Hosting Autopsy

Autopsy

The auto-reply that answered itself eleven thousand times

A company set up an out-of-office reply on a shared mailbox. Separately, a helpdesk system sent an automatic...

Autopsy

The one-line .htaccess edit that took three sites down

An administrator added a redirect rule to a shared .htaccess file in a parent folder, which three sites...

Autopsy

The firewall rule that blocked the payment provider

After a burst of suspicious traffic, a developer added a rule that blocked all requests from addresses...