The Host's Casebook / The agency and the staging site that got indexed

The agency and the staging site that got indexed

CASEBOOK

6 min read · 1,297 words

A note on authenticity. This is a composite story written by the editors, built from situations that come up again and again. It is not the account of a particular named person or business. Real reader stories go through the submission page and are marked as reader-submitted.

A small web agency built a new website for a client, a regional furniture retailer, on a staging subdomain. Staging is the normal practice: a private copy of the site where the client can review pages, the developers can break things, and nobody outside the project sees any of it. When the client approved the design, the agency pushed the site live on the main domain and moved on to the next job. The staging copy stayed where it was, still running, because nobody had a task on the board that said to take it down.

A few weeks later the client searched for one of their own product categories and saw two results that looked almost identical. One was the new site. The other was the same page from the staging address, wearing the old design that had been replaced. When the client checked others, the pattern repeated: two versions of nearly every page.

What had been left behind

The staging site lived at an address like staging.example.com. It was a full copy: the same pages, the same product listings, the same images, a database snapshot from the point of the last sync. It had no password, because early in the project the agency had wanted the client to review it without a login getting in the way. It was not linked from anywhere obvious.

It did, however, have one protection that had been working: WordPress has a setting under Reading that asks search engines not to index the site. While it is on, the site adds a "noindex" instruction to its pages, and well-behaved crawlers leave it alone. A developer had switched it off a few weeks earlier to test an SEO plugin's sitemap output, meaning to turn it back on, and then had been pulled onto something else.

That setting is a request, not a lock. Search engines normally honour it, but anything else can still read the site, and a request is only as good as the last person who remembered to leave it switched on.

How the search engines found it

Nobody had submitted the staging address anywhere. Crawlers do not need an invitation; they follow links. In this case a single link was enough. An image on the live homepage still pointed at a file on the staging subdomain, because the content had been copied across with absolute URLs. A crawler reading the live site saw that image URL, noted the hostname, and requested the root of that hostname to see what it was. With no password and no noindex, it walked in.

Searchcrawlerwww.example.comlive sitestaging.example.comno password, no noindexSearch indexboth copiesone image link
One leftover link from the live site was enough for a crawler to discover and index the whole staging copy.

Once the staging pages were in the index, the search engine had two near-identical sets of pages to choose between. It does its best to pick a canonical version, but without clear signals it can pick wrongly or show both. Some product pages from staging outranked the live ones for a while, because they had older, and in some cases stronger, history.

What the client actually saw

The symptoms appeared in an order that looks obvious in hindsight and was confusing at the time.

The client's first reaction was to blame the new site. The agency's first reaction was to check the live pages, which were fine. It took a day before someone looked at the address in the search result rather than the page.

The fix, in the order that works

Cleaning this up has a correct order, and doing it out of order can make things last longer.

  1. Put a password on staging first. The agency added HTTP basic authentication at the web server, which stops crawlers and humans before PHP or WordPress is involved. In Apache this is a few lines in the virtual host or an .htaccess file:
    AuthType Basic
    AuthName "Staging"
    AuthUserFile /home/agency/.htpasswd
    Require valid-user
  2. Turn the noindex back on in WordPress, and also send it as a header from the server, so it survives a database copy: Header set X-Robots-Tag "noindex, nofollow".
  3. Ask for removal. In the search engine's webmaster tool, they used the removals feature on the staging property, which hides the host from results fairly quickly but only for a limited time, around six months.
  4. Wait for the crawler to notice. The old entries only drop out once the crawler has come back and been met with a 401 or a noindex, which is outside anyone's control.
  5. Fix the cause on the live site. They searched the database for the staging hostname and replaced the absolute URLs, so that no live page linked to staging again.
  6. Add canonical tags on the live pages pointing at themselves, which gives the search engine a clear statement of the preferred version.
Do not block staging in robots.txt and call it done. A crawler that is told not to fetch a page cannot see its noindex tag, and a blocked address can still appear in results as a bare link. Use a password; it is the only option that cannot be read the wrong way.

Why it took a couple of months

The agency expected the duplicates to vanish in days. They took about two months. Search engines revisit pages on their own schedule, which for a small retailer's lower-priority pages can be weeks apart. Each staging page had to be fetched again, seen to be gone or blocked, and dropped from the index. The removal tool hid results quickly, but the underlying index entries took their own time.

Day 0Password addedDay 2Removal requestWeeks 3-5Crawler revisits, drops pagesAbout week 8Duplicates goneIllustrative timings for a small site
The quick actions happen in days; the index catches up over weeks.

A launch checklist that would have prevented it

The agency wrote a short list for the day a site goes live, and it is worth copying. None of the items takes more than a few minutes, and the whole list is quicker than the clean-up was.

Commands worth running

Search for your own site with the operator site:example.com and look at which hostnames appear. Anything that is not your public site is a leak. Then test each non-public hostname from the command line:

curl -I https://staging.example.com/
curl -s https://staging.example.com/ | grep -i robots

A protected site returns 401 Unauthorized on the first request. A site that returns 200 and shows no robots meta tag or X-Robots-Tag header is open to the world. Also list the subdomains you have in DNS, since old test hostnames are easy to forget. The DNS cheatsheet has the commands, and the redirect generator helps when you want a retired staging address to send visitors to the live site instead.

PreviousThe club newsletter that went to spamNextThe photographer and the 4 MB hero image

More from The Host's Casebook

Composite case

The podcast whose episodes lived on a free plan

A hobby podcaster hosted audio files on a free website plan to avoid paying for podcast hosting. It worked...

Composite case

The tutor who wanted a proper email address

A private tutor had a free webmail address on her business cards and wanted something with her own name. She...

Composite case

The hobby forum and the registration flood

A small forum for vintage radio enthusiasts woke up to find 4,000 new members overnight. None had posted yet....