When an online business moved from one server to two, to be safe, it copied the whole server image. That image included the crontab, with a nightly job that generated and emailed invoices.
That night both servers ran the job. Every customer received two invoices, and the accounting system created two ledger entries for each. Sorting it out took a bookkeeper most of a week.
They moved scheduled work to a single designated machine and added a lock so a job refuses to run if another copy is active.
Why they went from one server to two
The business sold subscription boxes online and had grown from a few hundred customers to a few thousand. Its single server was coping, but just. Every month-end the checkout slowed, and a failed disk a year earlier had left the owner nervous about having everything on one machine.
The plan was reasonable: add a second server so that traffic could be shared and one machine could fail without taking the shop down. The hosting provider offered a way to clone a server image, and the developer used it. A clone is attractive because it is fast and exact. Everything that worked on the first server works on the second, including the parts nobody remembers installing.
That is also the trap. "Everything" includes the scheduled jobs, the mail settings and the integration keys, and some of those are meant to exist exactly once.
What ran that night
The crontab held three lines, written by a previous contractor: a database backup at 02:00, a stock-sync with the warehouse at 01:00, and the invoice job at 01:30. After the clone, the second server had identical copies of all three. At 01:30 both machines woke up, queried the same database for the day's completed orders, rendered an invoice PDF for each and sent it by email.
The two runs were milliseconds apart in their start but not in their results. Each saw the same set of orders marked "invoice not yet sent", because neither had finished and updated the flag when the other began reading. That gap is called a race, and it is the classic way two copies of a job produce duplicates.
Symptoms, as people noticed them
The first sign came at 07:40, when a customer replied to what she took to be a billing error. By 09:00 the support inbox held thirty similar messages, and the owner assumed the email system had a fault. The accountant noticed separately that the ledger's morning import had twice as many lines as the day before.
Duplicate invoices are a reputational problem, since customers wonder whether they will be charged twice. They are also a bookkeeping one, because each invoice creates a receivable and each receivable has to be reversed with a credit note. The bookkeeper matched, voided and annotated around two thousand of them. Nothing was charged twice, because payments ran from a different job, but explaining that to customers one by one took longer than the correction.
Wrong turns
The first theory was a bug in the invoice script, so the developer read it carefully and found nothing wrong. The second was a mail provider retry, which the provider's log disproved: each message had been submitted twice, from two different server addresses. That detail was the giveaway. The email headers showed two origin hosts, and only one of them was supposed to be sending mail.
Once the developer logged in to the second machine and ran crontab -l, the cause was on the screen in three lines.
The fix: one owner, plus a lock
The immediate step was to comment out the three jobs on the second server. The lasting one was to make a deliberate choice about which machine owns scheduled work, and to write it down. Many teams call this the "cron box" or the scheduler host. It runs jobs; the others serve web traffic.
The second protection is a lock, so a job refuses to start if another copy is already running, on the same machine or a different one. On a single machine, flock does it with one line:
30 1 * * * flock -n /var/lock/invoices.lock php /srv/shop/scripts/invoices.php
The -n makes the second attempt give up instead of waiting. Across two machines, a file lock on local disk is no use, because each machine has its own. There you need a lock both can see, usually a row in the shared database that the job claims at the start.
Making the job safe to run twice
Locks are good, but better is a job that cannot do harm when it runs twice. The word for that is idempotent: running it again produces the same result as running it once. For invoices, the script now sets a flag on each order inside a single database transaction before sending, so a second reader sees "claimed" and skips it. Each invoice also carries a unique number tied to the order, and the ledger rejects a number it has already booked.
Both together mean the fault would need to survive two independent defences to hurt anyone.
Ways to make a job run once
| Approach | Works across machines | Catch |
|---|---|---|
| Jobs only in one crontab | Yes, by convention | Depends on people remembering; a clone breaks it |
flock on a local file | No | Protects against overlap on one machine only |
| Lock row in shared database | Yes | Needs a timeout so a crashed job does not block for ever |
| Idempotent job design | Yes | Takes thought per job, and testing |
| Hosting platform scheduler | Yes | Tied to the provider's tooling |
The timeout in row three is the detail people forget. If the job crashes halfway, its lock row stays, and the next night's run refuses to start because it believes someone else is working. Give the lock an expiry a little longer than the job's normal runtime.
The week after
The apology email went out the same afternoon, two sentences long, saying that duplicate invoices had been sent in error, that no one had been charged twice and that no action was needed. It cut the support inbox's volume by about half within an hour. The remaining customers wanted to be told, individually, that their own account was fine.
The bookkeeper's week was the real bill. Each duplicate needed a credit note, and each credit note needed a reference to the original, which the system could not generate in bulk. The owner later reckoned the clone had saved a day of setup and cost five days of cleanup, which is not a trade anyone would choose knowingly.
One useful side effect: the team finally wrote down every scheduled job in a shared document, with its purpose, its owner and the machine it belongs on. The document has six entries. Before the incident nobody could have listed more than three.
A ten-minute check
On any server cloned from another, run crontab -l for every user that has one, look in /etc/cron.d, and list timers with systemctl list-timers. Compare the result with the other servers and decide, for each job, whether it should run on every machine (log cleanup) or once in total (invoices, reports, newsletters, anything that sends or charges). Also look at the web application's own scheduler: WordPress's page-triggered cron, for instance, runs on every web server that receives traffic.
The cron helper turns schedules into readable English, which makes a review quicker.
Things people ask
Is a load-balanced setup always like this?
Not always, but anything that scales out has to answer the question of who runs the timers. Many managed platforms answer it for you; ask.
Could this happen on shared hosting?
Rarely, but a staging copy that includes the production crontab and the live mail settings is the same mistake in a smaller box.
Why not just clone and then edit?
You can, and should. The mistake was treating the clone as finished.
What about the stock sync and the backup, which also ran twice?
They did, and were harmless by luck. The sync wrote the same figures twice, and the backup produced two files with different names. Neither would stay harmless for long: two simultaneous backups can slow a database to a crawl, and two syncs with a changing warehouse feed can overwrite each other with stale numbers.
What would have caught it
- Decide which machine owns scheduled work when you scale out.
- Make jobs idempotent so running twice is harmless, or add a lock.
- Review the crontab on any cloned server.
- Check the headers of one real email after the change, to see which machine sent it.