Hosting Autopsy / The sitemap that listed forty thousand junk URLs

The sitemap that listed forty thousand junk URLs

HOSTING AUTOPSY

7 min read · 1,535 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A furniture retailer added a filter system to its catalogue: colour, material, price band, size. Each combination produced a unique address, and a plugin dutifully added every one of them to the XML sitemap. A catalogue of 600 products turned into a sitemap of about 40,000 near-identical pages.

Search crawlers spent their visits on filter combinations nobody would ever search for, and the product pages that mattered were crawled less often. Rankings slid slowly enough that nobody connected it to the filter launch three months earlier. The hosting bill rose too, because the crawlers were generating real server load.

The fix was to exclude filter addresses from the sitemap, mark them noindex, and block the heaviest patterns in robots.txt. Recovery took a couple of months. This is a composite of several shops with the same problem, and the numbers below are illustrative, but the arithmetic of how a filter menu explodes is real and worth seeing laid out.

How 600 products became 40,000 addresses

The retailer sold sofas, tables, shelving and beds. Customers had asked for filters, so the developer added four: colour, material, price band and size. Each filter was a query parameter, so a visitor choosing oak and then white ended up on an address like /tables/?material=oak&colour=white.

The trouble is that filters multiply. Suppose a category has 8 colours, 5 materials, 6 price bands and 5 sizes. Each filter can also be left unset, so the choices are 9 x 6 x 7 x 6 combinations, which is 2,268 for that one category. Add the order of parameters (colour first or material first gives a different address with the same content) and sorting options, and one category yields tens of thousands of distinct addresses.

Across a catalogue of a dozen categories, the total passed 40,000. Almost all of them listed the same few dozen products, shuffled and trimmed. To a person, that is a helpful menu. To a crawler that reads addresses, it is an enormous, repetitive website.

Addresses in the sitemap (illustrative)Before filtersabout 700After filtersabout 40,000Bar lengths are in proportion. Product and category pages did not change.
The pages worth finding stayed the same; the list of addresses grew by a factor of about fifty.
Where crawler requests went (illustrative, share of a day)Filter launch15%filter pages 85%After the fixproduct and category pages 93%Green: pages worth finding. Amber: filter combinations.
The same crawler visits, spent very differently before and after the exclusion.

What crawlers do with a sitemap

A sitemap is a list of addresses you would like looked at, with optional hints about when each last changed. Crawlers treat it as a suggestion and not a command, and they budget their visits to a site according to how quickly it answers and how useful the earlier pages proved. A site that offers 40,000 near-duplicates teaches the crawler that most addresses are not worth the trip, so it visits each less often, including the good ones.

That is why the harm was indirect. Nothing was removed or penalised. The valuable pages were merely visited later, and a page that is visited later is updated later in search results.

What the plugin did, and why nobody noticed

The sitemap was generated by a plugin that walked the shop's database and listed every public address it could produce. When the filter feature arrived, the plugin treated each filter result as a normal listing page and added it. Nobody had told it otherwise, and nobody looked at the file afterwards. A sitemap is a file most people never open once it exists.

There was also nothing to see on the site. Visitors browsed as before. The shop's own analytics showed normal traffic. The only people who experienced the 40,000 addresses were robots, and robots do not file support tickets.

Meanwhile the sitemap was being submitted to the search consoles automatically on every change. The files were large enough to be split into several parts of up to 50,000 addresses each, which is the standard limit for a single sitemap file, so even the splitting looked orderly.

Symptoms, in the order they appeared

The first sign was in the server logs, though nobody was reading them with that in mind. Within weeks of the launch, a significant share of requests came from well-known crawlers and almost all of them were for filter addresses. Pages built by combining four filters are slow, because each one runs a database query with several conditions, and the crawlers asked for thousands in a row.

The second sign was the hosting bill. The shop was on a managed plan with resource limits, and it began to hit them during the early hours, when crawlers were most active. An upgrade to the next plan was suggested and accepted, which treated the symptom.

The third sign took longer: the product pages that earned the money began to slip. New products took weeks to appear in search, where they used to take days. Older pages were refreshed less often, so price changes and stock changes lagged behind reality. By the time the owner noticed the slide in rankings, three months had gone by, and the filter launch was long forgotten as a cause.

Wrong turns

The first theory was seasonal. Furniture sales dip in some months, and the owner assumed the search slide was the same pattern. The second was content: the shop rewrote a number of product descriptions, which cost money and changed nothing.

Someone then suspected a penalty, and spent a week reading guidance on link quality. A different consultant proposed buying links. Fortunately nobody did.

The useful clue came from the search console's crawl statistics report. It showed the number of requests per day roughly quadrupling around the date of the filter launch, while the share of requests for product pages fell sharply. Alongside it, the coverage report listed tens of thousands of addresses as "discovered" or "crawled but not indexed". That is what a crawler's frustration looks like on a chart.

The fix, step by step

The repair had three parts and the order mattered.

  1. Take filter addresses out of the sitemap. The plugin had a setting to exclude addresses carrying query parameters, which nobody had turned on. Once it was, the sitemap fell back to about 700 entries: products, categories, and a few information pages.
  2. Mark the filter pages noindex, follow, so a crawler that arrives by a link understands they are not meant for search results. The filter pages also gained a canonical link pointing at the unfiltered category.
  3. Block the heaviest patterns in robots.txt, such as sort orders and price bands, to take load off the server.
User-agent: *
Disallow: /*?*sort=
Disallow: /*?*price=
Sitemap: https://www.example.com/sitemap.xml
Do not block in robots.txt a page you have also marked noindex. A crawler that is forbidden to fetch the page never sees the noindex instruction, so already-indexed addresses can linger. Apply noindex first, wait for the pages to drop out, and only then block the patterns.

The retailer got the order slightly wrong at first, and the clean-up took longer for it.

Recovery and aftermath

Crawl statistics began to normalise within a fortnight of the sitemap change. The server load fell back and the shop dropped to its previous hosting plan a month later. Rankings took longer, since search engines revisit pages at their own pace. A couple of months passed before the product pages were back where they had been.

The shop added two habits. After any feature launch, someone opens the sitemap and counts. And the developer wrote a one-page note for the next person: filters are for visitors, not for search results, and any plugin that lists addresses should be tested on a staging copy first.

A ten-minute check

Open /sitemap.xml in a browser and scroll. If it is an index, open each part. Then count entries from the command line:

curl -s https://www.example.com/sitemap-products.xml | grep -c "<loc>"
curl -s https://www.example.com/sitemap-products.xml | grep "<loc>" | grep -c "?"

The first number is how many addresses you are advertising. The second is how many carry a query string. For a shop, the second should be at or near zero. Then check your own logs for crawler activity on parameter addresses:

grep -i "bot" access.log | grep -c "?colour="

The status code reference helps with reading the responses, and the troubleshooting guide covers load problems more broadly.

Loose ends

Should filter pages ever be indexed?

Occasionally, yes. A page for "oak dining tables" may match something people really search for. The usual approach is to pick a few such combinations on purpose, give them clean addresses and titles, and keep the rest out.

Does a bigger sitemap mean more visits from crawlers?

It means more requests, not better ones. A sitemap is a suggestion about what matters, so padding it dilutes the suggestion.

Is robots.txt enough on its own?

No. It controls fetching, not indexing, as the note above explains.

How often should the sitemap be reviewed?

After every feature launch, and otherwise a few times a year.

What would have caught it

PreviousThe disk that filled up overnightNextThe PHP upgrade that broke the checkout

More from Hosting Autopsy

Autopsy

The disk that filled up overnight

A VPS hosted four or five small sites without problems for two years. One morning all of them returned...

Autopsy

The one-line .htaccess edit that took three sites down

An administrator added a redirect rule to a shared .htaccess file in a parent folder, which three sites...

Autopsy

The migration that forgot the cron jobs

A small publishing company moved to a new host over a weekend. The work was done carefully: files copied...