A furniture retailer added a filter system to its catalogue: colour, material, price band, size. Each combination produced a unique address, and a plugin dutifully added every one of them to the XML sitemap. A catalogue of 600 products turned into a sitemap of about 40,000 near-identical pages.
Search crawlers spent their visits on filter combinations nobody would ever search for, and the product pages that mattered were crawled less often. Rankings slid slowly enough that nobody connected it to the filter launch three months earlier. The hosting bill rose too, because the crawlers were generating real server load.
The fix was to exclude filter addresses from the sitemap, mark them noindex, and block the heaviest patterns in robots.txt. Recovery took a couple of months. This is a composite of several shops with the same problem, and the numbers below are illustrative, but the arithmetic of how a filter menu explodes is real and worth seeing laid out.
How 600 products became 40,000 addresses
The retailer sold sofas, tables, shelving and beds. Customers had asked for filters, so the developer added four: colour, material, price band and size. Each filter was a query parameter, so a visitor choosing oak and then white ended up on an address like /tables/?material=oak&colour=white.
The trouble is that filters multiply. Suppose a category has 8 colours, 5 materials, 6 price bands and 5 sizes. Each filter can also be left unset, so the choices are 9 x 6 x 7 x 6 combinations, which is 2,268 for that one category. Add the order of parameters (colour first or material first gives a different address with the same content) and sorting options, and one category yields tens of thousands of distinct addresses.
Across a catalogue of a dozen categories, the total passed 40,000. Almost all of them listed the same few dozen products, shuffled and trimmed. To a person, that is a helpful menu. To a crawler that reads addresses, it is an enormous, repetitive website.
What crawlers do with a sitemap
A sitemap is a list of addresses you would like looked at, with optional hints about when each last changed. Crawlers treat it as a suggestion and not a command, and they budget their visits to a site according to how quickly it answers and how useful the earlier pages proved. A site that offers 40,000 near-duplicates teaches the crawler that most addresses are not worth the trip, so it visits each less often, including the good ones.
That is why the harm was indirect. Nothing was removed or penalised. The valuable pages were merely visited later, and a page that is visited later is updated later in search results.
What the plugin did, and why nobody noticed
The sitemap was generated by a plugin that walked the shop's database and listed every public address it could produce. When the filter feature arrived, the plugin treated each filter result as a normal listing page and added it. Nobody had told it otherwise, and nobody looked at the file afterwards. A sitemap is a file most people never open once it exists.
There was also nothing to see on the site. Visitors browsed as before. The shop's own analytics showed normal traffic. The only people who experienced the 40,000 addresses were robots, and robots do not file support tickets.
Meanwhile the sitemap was being submitted to the search consoles automatically on every change. The files were large enough to be split into several parts of up to 50,000 addresses each, which is the standard limit for a single sitemap file, so even the splitting looked orderly.
Symptoms, in the order they appeared
The first sign was in the server logs, though nobody was reading them with that in mind. Within weeks of the launch, a significant share of requests came from well-known crawlers and almost all of them were for filter addresses. Pages built by combining four filters are slow, because each one runs a database query with several conditions, and the crawlers asked for thousands in a row.
The second sign was the hosting bill. The shop was on a managed plan with resource limits, and it began to hit them during the early hours, when crawlers were most active. An upgrade to the next plan was suggested and accepted, which treated the symptom.
The third sign took longer: the product pages that earned the money began to slip. New products took weeks to appear in search, where they used to take days. Older pages were refreshed less often, so price changes and stock changes lagged behind reality. By the time the owner noticed the slide in rankings, three months had gone by, and the filter launch was long forgotten as a cause.
Wrong turns
The first theory was seasonal. Furniture sales dip in some months, and the owner assumed the search slide was the same pattern. The second was content: the shop rewrote a number of product descriptions, which cost money and changed nothing.
Someone then suspected a penalty, and spent a week reading guidance on link quality. A different consultant proposed buying links. Fortunately nobody did.
The useful clue came from the search console's crawl statistics report. It showed the number of requests per day roughly quadrupling around the date of the filter launch, while the share of requests for product pages fell sharply. Alongside it, the coverage report listed tens of thousands of addresses as "discovered" or "crawled but not indexed". That is what a crawler's frustration looks like on a chart.
The fix, step by step
The repair had three parts and the order mattered.
- Take filter addresses out of the sitemap. The plugin had a setting to exclude addresses carrying query parameters, which nobody had turned on. Once it was, the sitemap fell back to about 700 entries: products, categories, and a few information pages.
- Mark the filter pages
noindex, follow, so a crawler that arrives by a link understands they are not meant for search results. The filter pages also gained a canonical link pointing at the unfiltered category. - Block the heaviest patterns in
robots.txt, such as sort orders and price bands, to take load off the server.
User-agent: *
Disallow: /*?*sort=
Disallow: /*?*price=
Sitemap: https://www.example.com/sitemap.xml
The retailer got the order slightly wrong at first, and the clean-up took longer for it.
Recovery and aftermath
Crawl statistics began to normalise within a fortnight of the sitemap change. The server load fell back and the shop dropped to its previous hosting plan a month later. Rankings took longer, since search engines revisit pages at their own pace. A couple of months passed before the product pages were back where they had been.
The shop added two habits. After any feature launch, someone opens the sitemap and counts. And the developer wrote a one-page note for the next person: filters are for visitors, not for search results, and any plugin that lists addresses should be tested on a staging copy first.
A ten-minute check
Open /sitemap.xml in a browser and scroll. If it is an index, open each part. Then count entries from the command line:
curl -s https://www.example.com/sitemap-products.xml | grep -c "<loc>"
curl -s https://www.example.com/sitemap-products.xml | grep "<loc>" | grep -c "?"
The first number is how many addresses you are advertising. The second is how many carry a query string. For a shop, the second should be at or near zero. Then check your own logs for crawler activity on parameter addresses:
grep -i "bot" access.log | grep -c "?colour="
The status code reference helps with reading the responses, and the troubleshooting guide covers load problems more broadly.
Loose ends
Should filter pages ever be indexed?
Occasionally, yes. A page for "oak dining tables" may match something people really search for. The usual approach is to pick a few such combinations on purpose, give them clean addresses and titles, and keep the rest out.
Does a bigger sitemap mean more visits from crawlers?
It means more requests, not better ones. A sitemap is a suggestion about what matters, so padding it dilutes the suggestion.
Is robots.txt enough on its own?
No. It controls fetching, not indexing, as the note above explains.
How often should the sitemap be reviewed?
After every feature launch, and otherwise a few times a year.
What would have caught it
- Look at what your sitemap actually contains after every feature launch.
- Compare the number of sitemap URLs with the number of pages you would want found.
- Watch crawl statistics in the search console for sudden growth.
- Count requests from crawlers on addresses with query strings in the server log every month.
- Test any plugin that generates addresses on a staging copy and read its output before it goes live.