A small web agency built a new website for a client, a regional furniture retailer, on a staging subdomain. Staging is the normal practice: a private copy of the site where the client can review pages, the developers can break things, and nobody outside the project sees any of it. When the client approved the design, the agency pushed the site live on the main domain and moved on to the next job. The staging copy stayed where it was, still running, because nobody had a task on the board that said to take it down.
A few weeks later the client searched for one of their own product categories and saw two results that looked almost identical. One was the new site. The other was the same page from the staging address, wearing the old design that had been replaced. When the client checked others, the pattern repeated: two versions of nearly every page.
What had been left behind
The staging site lived at an address like staging.example.com. It was a full copy: the same pages, the same product listings, the same images, a database snapshot from the point of the last sync. It had no password, because early in the project the agency had wanted the client to review it without a login getting in the way. It was not linked from anywhere obvious.
It did, however, have one protection that had been working: WordPress has a setting under Reading that asks search engines not to index the site. While it is on, the site adds a "noindex" instruction to its pages, and well-behaved crawlers leave it alone. A developer had switched it off a few weeks earlier to test an SEO plugin's sitemap output, meaning to turn it back on, and then had been pulled onto something else.
How the search engines found it
Nobody had submitted the staging address anywhere. Crawlers do not need an invitation; they follow links. In this case a single link was enough. An image on the live homepage still pointed at a file on the staging subdomain, because the content had been copied across with absolute URLs. A crawler reading the live site saw that image URL, noted the hostname, and requested the root of that hostname to see what it was. With no password and no noindex, it walked in.
Once the staging pages were in the index, the search engine had two near-identical sets of pages to choose between. It does its best to pick a canonical version, but without clear signals it can pick wrongly or show both. Some product pages from staging outranked the live ones for a while, because they had older, and in some cases stronger, history.
What the client actually saw
The symptoms appeared in an order that looks obvious in hindsight and was confusing at the time.
- Search results showed two entries per page, with different address bars and one in the old design.
- A few customers phoned, confused, because a product on the staging copy had an old price and could not be bought, since the checkout there was switched off.
- Organic traffic to the live site dipped in a way that was hard to pin on any other cause.
- In the search tool's reports, the live site's category pages showed up as "duplicate" or "alternate page with proper canonical" entries, and the staging host appeared among the indexed addresses.
The client's first reaction was to blame the new site. The agency's first reaction was to check the live pages, which were fine. It took a day before someone looked at the address in the search result rather than the page.
The fix, in the order that works
Cleaning this up has a correct order, and doing it out of order can make things last longer.
- Put a password on staging first. The agency added HTTP basic authentication at the web server, which stops crawlers and humans before PHP or WordPress is involved. In Apache this is a few lines in the virtual host or an
.htaccessfile:AuthType Basic AuthName "Staging" AuthUserFile /home/agency/.htpasswd Require valid-user - Turn the noindex back on in WordPress, and also send it as a header from the server, so it survives a database copy:
Header set X-Robots-Tag "noindex, nofollow". - Ask for removal. In the search engine's webmaster tool, they used the removals feature on the staging property, which hides the host from results fairly quickly but only for a limited time, around six months.
- Wait for the crawler to notice. The old entries only drop out once the crawler has come back and been met with a 401 or a noindex, which is outside anyone's control.
- Fix the cause on the live site. They searched the database for the staging hostname and replaced the absolute URLs, so that no live page linked to staging again.
- Add canonical tags on the live pages pointing at themselves, which gives the search engine a clear statement of the preferred version.
robots.txt and call it done. A crawler that is told not to fetch a page cannot see its noindex tag, and a blocked address can still appear in results as a bare link. Use a password; it is the only option that cannot be read the wrong way.Why it took a couple of months
The agency expected the duplicates to vanish in days. They took about two months. Search engines revisit pages on their own schedule, which for a small retailer's lower-priority pages can be weeks apart. Each staging page had to be fetched again, seen to be gone or blocked, and dropped from the index. The removal tool hid results quickly, but the underlying index entries took their own time.
A launch checklist that would have prevented it
The agency wrote a short list for the day a site goes live, and it is worth copying. None of the items takes more than a few minutes, and the whole list is quicker than the clean-up was.
- Search the live database and theme files for the staging hostname, and replace any absolute links to it.
- Confirm the live site's search-visibility setting is on, and staging's is off, since the two are easy to swap after copying a database across.
- Decide what happens to the staging copy: lock it with a password, or delete it, and record which in the project notes.
- Add the live site to the search engine's webmaster tool and check the sitemap lists only live addresses.
- Look again two weeks after launch with a
site:search, because that is when leaks like this one start to show.
Commands worth running
Search for your own site with the operator site:example.com and look at which hostnames appear. Anything that is not your public site is a leak. Then test each non-public hostname from the command line:
curl -I https://staging.example.com/
curl -s https://staging.example.com/ | grep -i robots
A protected site returns 401 Unauthorized on the first request. A site that returns 200 and shows no robots meta tag or X-Robots-Tag header is open to the world. Also list the subdomains you have in DNS, since old test hostnames are easy to forget. The DNS cheatsheet has the commands, and the redirect generator helps when you want a retired staging address to send visitors to the live site instead.