Hosting Autopsy / The autoscaler that scaled to meet the bots

The autoscaler that scaled to meet the bots

HOSTING AUTOPSY

6 min read · 1,349 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A small software startup ran its public site and product pages on a platform with auto-scaling, because the founders had been burned before by a launch-day outage and never wanted that again. The rule was simple: when average CPU across the web servers passes 60 per cent, add another server, up to whatever the account allows. When load falls, remove them again.

For eight months it did what it said. Then, one weekend, a scraper began requesting every page of the site, then every page again with different query strings, as fast as it could. The platform saw the load and added servers to meet it. It kept adding them until Monday.

The site stayed up the entire time. The invoice at the end of the month was nearly ten times the usual amount, and every extra pound of it had been spent serving one scraper. This is a composite, with illustrative figures, but the shape of it is common enough that most hosting engineers have a version.

What auto-scaling actually does

Auto-scaling is a feedback loop with a price tag attached. It watches a metric, usually CPU, memory or requests per instance, and compares it with a target. If the metric is above the target, it starts more instances behind the load balancer. If it is below the target for long enough, it stops some.

What it cannot do is decide whether the load is worth serving. A real crowd arriving from a news mention and a script fetching the same catalogue page forty times a second look identical at the CPU graph. The platform is doing precisely what it was configured to do, which is the problem: it is a loyal servant with no opinion about whether the request is a good one.

Real visitors Scraper Load balancer Autoscaling group server 1 server 2 server 3 (added) server 4... (added)
Nothing between the internet and the scaling group can tell a customer from a script, so both get capacity.

The weekend, hour by hour

The scraper started on Friday evening. It was polite for the first hour, a few requests a second, and then it opened more connections. By midnight the group had gone from its normal two servers to six. By Saturday midday it was at fourteen, and by Saturday evening it had reached the account's limit of twenty, which nobody had ever thought to lower.

Because the extra servers kept CPU under the threshold, the site stayed fast. Monitoring showed green. The uptime checker reported 100 per cent for the weekend. The only red thing anywhere was a billing dashboard that nobody was looking at.

The first human clue came on Monday when an engineer glanced at the instance list and saw twenty servers where there should have been two. Even then the first reaction was relief that they had coped so well with a traffic spike, until someone looked at the analytics and found no matching rise in real visitors.

Illustrative: servers (warm line) and real visitors (flat line) FriSatSun 20 servers, the account limit real visitors, unchanged
Capacity climbs to the ceiling while the number of genuine visitors does not move.

Where the money went

The extra bill had three parts. The biggest was compute: eighteen additional servers for roughly sixty hours. The second was outbound data transfer, because the scraper downloaded product images at full size. The third was a managed database that scaled up its connection tier when the web servers multiplied.

ItemNormal month (illustrative)This month (illustrative)
Web servers£180£1,650
Data transfer out£40£780
Database tier£90£470
Total£310about £2,900

The numbers are small by company standards and large by startup standards, and the near ten-fold multiplier is the thing to take away. The sums are illustrative, but the proportions are typical: transfer and database costs grow alongside compute, so watching the server line alone understates the damage. A flat-priced shared or VPS plan would have slowed down and perhaps fallen over, which is a different kind of bad, but it would not have produced an invoice.

Diagnosis and fix

The access log made the cause obvious. A quick count of requests by client address, then by user agent, showed three address ranges from one hosting provider making over 90 per cent of the traffic, all with the same generic library user agent.

awk '{print $1}' access.log | sort | uniq -c | sort -rn | head
  412903 198.51.100.17
  388211 198.51.100.18
  371008 198.51.100.19
    9120 203.0.113.55

The repairs were in layers. First, they blocked the three ranges at the edge, which brought load down at once. Second, they added rate limiting in front of the application so that no single address could make more than a set number of requests per minute. Third, they put a bot filter on the product and search paths, the favourite targets. Fourth, and the one they should have had from the start, they set a hard maximum of four servers and an alert at three.

A cache for the catalogue pages helped too. Most scraper requests were for pages that never change between deployments, and serving them from cache cost almost nothing.

Why a ceiling matters more than a floor

People tune the minimum carefully because it protects uptime. The maximum is where the money is. A maximum is a statement about how much you are prepared to pay to keep serving traffic at its worst, and it should be a deliberate number, picked by someone with budget authority, not a platform default.

The right ceiling is a little above the largest genuine peak you have measured, plus margin. If the ceiling is hit, you want to know about it as an event, not discover it on an invoice. An alert at 75 per cent of the ceiling is worth more than a dashboard.

The assumptions behind it

Looking back, the team had made three assumptions that were never written down. The first was that all traffic is customer traffic, so serving more of it is always good. The second was that a cap was unnecessary because they would "see it coming". The third was that cost and reliability are separate topics owned by different people: engineers watched the graphs, one founder paid the card bill, and nobody watched both.

None of these is unreasonable for a team of six. They are the normal result of setting something up in a hurry before a launch and then having it work. A system that works stops being looked at, which is why it deserves a calendar reminder to be looked at on purpose.

They also found, while reading the platform documentation properly for the first time, that the scale-in rule was slow by design. After the scraper was blocked on Monday, the servers took close to an hour to drain away, because the group waits for CPU to stay low through a cool-down period before removing anything. That is a fair protection against flapping, and it meant the meter was still running after the problem was solved.

A ten-minute check

Open the scaling configuration and read three numbers: minimum, maximum and the metric threshold. If the maximum is the platform default or blank, set one today. Then look at the billing page for a budget alert that emails a person who will read it, and set one at, say, 120 per cent of a typical month.

Finally, ask what is in front of the scaler. If the answer is "nothing", the load balancer is the only gate, and anyone on the internet can spend your money. Check the uptime maths tool if you are weighing how much capacity redundancy is really worth, and compare against page weight on the page weight checker, because lighter pages cost less to serve.

What would have caught it

PreviousThe image that cost a month's hostingNextThe HSTS setting that broke the intranet

More from Hosting Autopsy

Autopsy

The image that cost a month's hosting

A photographer on a pay-by-usage cloud plan published a picture that was picked up by a popular forum. People...

Autopsy

The IPv6 record pointing at nothing

After moving to a new host, a nonprofit updated the A record and thought no more of it. An old AAAA record...

Autopsy

The backup that could not be restored

A hobbyist forum had a nightly backup running faithfully for over a year. Every morning the control panel...