Hosting Autopsy / The robots.txt that told everyone to go away

The robots.txt that told everyone to go away

HOSTING AUTOPSY

7 min read · 1,432 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

A regional estate agency commissioned a new website from a freelance developer. While building it on a temporary address, the developer did what careful developers do: she did not want half-finished pages appearing in search results. So she ticked the WordPress option that asks search engines not to index the site, and she also added a robots.txt file that disallowed everything.

When the site went live, the files were copied across with the rest. Nobody remembered either of them, because both were invisible. A page with a robots rule looks exactly like a page without one.

Two switches that look like one

The block came from two places, and that is part of why it survived. The first was the WordPress setting under Settings, Reading, labelled "Discourage search engines from indexing this site". When ticked, WordPress adds a noindex instruction to every page, in a meta tag in the HTML and in an HTTP header on some versions. The second was a hand-written file at the root of the site:

User-agent: *
Disallow: /

Those two lines mean: every crawler, stay out of everything. A well-behaved search engine reads the file before fetching anything, sees the rule, and does not request the pages.

The two mechanisms do different jobs, and the difference matters when you come to repair things. A robots rule controls crawling: whether the bot may fetch the page. A noindex instruction controls indexing: whether a page it has fetched may appear in results. They interact in an awkward way. If robots.txt forbids the fetch, the crawler never sees the page, so it never sees a noindex either, and a blocked address can in some cases still appear in results as a bare link with no description, based on links from elsewhere.

Crawlerwants a page Reads/robots.txt Disallow: / means stop herepage never fetched Allowed: fetch the pagethen obey any noindex tag The order is always the same: robots.txt first, page second.
A site-wide Disallow stops the crawler before it reaches the page, so nothing on the page can help.

What the owner noticed, and when

For the first days nothing seemed wrong. Visitors who typed the address or followed the old bookmarks reached the new site, the phone rang as usual, and the agency's staff looked at the site many times a day without discovering anything. The old site's pages were already in the index, and the search engines kept showing them for a while, with the old addresses leading to redirects or errors.

A week after launch, the owner searched for the company name and saw only old results, and one with a description reading "No information is available for this page." Analytics showed organic visits falling off a cliff: roughly 400 a day became about 60, almost all of them returning visitors who knew the name. The new site had been live and invisible for nine days.

The phrase that gave the game away appeared when the owner clicked that odd search result: the search engine could not show a description because it was not allowed to read the page. That is the visible symptom of a robots block, and it is easy to miss unless you know it.

Finding it

The first guess was a penalty for changing the site, which is a popular theory and rarely right. The second was that the new site was too slow. The developer, called back, thought of the staging settings in about thirty seconds, which says something about the value of asking the person who built it.

The check itself takes seconds:

curl -s https://example.com/robots.txt
curl -sI https://example.com/ | grep -i x-robots-tag
curl -s https://example.com/ | grep -i "name=.robots."

The first command showed the two lines above. The third showed <meta name='robots' content='noindex, nofollow' /> in the page head. Two blocks, both left from the build.

The search console, which the agency had not opened since the old site's days, would have said it plainly under its indexing reports: pages blocked by robots.txt, excluded by noindex tag. It also offers a tester that fetches a page the way the crawler would.

Removing the block, and why recovery took weeks

Removing the block was quick. The settings box was unticked and the robots file was replaced by one that allowed everything and pointed to the sitemap:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap.xml

Getting back to where they had been took several weeks. Search engines do not re-check a robots file on every request; they cache it for a while, often up to a day, and then they need to recrawl the pages, which they schedule by how important they judge the site to be. Several important pages, such as the area guides that had brought in most of the traffic, had been dropped from the index during the nine days and had to be found again from scratch. Submitting the sitemap and requesting indexing of the key pages in the search console sped it up, but did not make it instant.

LaunchBlock removed Daily organic visits over about eight weeks (illustrative)
The drop is immediate on the visible side, but the climb back follows the search engine's schedule, not the owner's.

Commands worth running

Make this a launch routine and repeat it after any migration:

  1. Open /robots.txt in a browser and read every line. If it says Disallow: / under User-agent: *, stop.
  2. In WordPress, look at Settings, Reading. The search visibility box must be unticked on the live site.
  3. View the page source of the homepage and search for noindex.
  4. Run curl -I on a few pages and look for an X-Robots-Tag header, which some servers and plugins add.
  5. Submit the sitemap in the search console and open the indexing report a day or two later.

Also protect the staging site properly. A password on the whole staging address (HTTP authentication) keeps out both crawlers and curious visitors, and it is not copied to the live site by accident the way a file is. For the status codes you may meet in these checks, see the status code reference.

MethodStops crawlingStops listingTravels to live by mistake
robots.txt DisallowYesNot reliablyYes, it is a file
Noindex tag or settingNoYesYes, if in the database
HTTP password on stagingYesYesNo, it belongs to the server

Quick answers

Is robots.txt a security measure?

No. It is a polite request, and the file itself is public. Never list private paths in it as if that hid them.

How long until the pages return?

For a small site, days to a few weeks. Important pages with many links tend to come back first.

Can I have an empty robots.txt?

Yes, an empty file means everything is allowed, and so does having none. A sitemap line is a useful addition.

Why does a blocked page still appear in results?

Because other sites link to it, the search engine may list the address without reading it. Allow crawling and use noindex if you want it gone.

What the agency changed afterwards

The agency agreed a rule with its developer: staging sites live behind a password, never behind a robots file alone. The launch checklist, now a shared document, has a line for the search visibility box, a line for robots.txt and a line for opening the homepage source and searching for noindex. Each line has a name and a date next to it.

They also learnt something about the timing of monitoring. The nine days were not a failure of any one person; they were a gap between the moment the site went live and the moment anybody looked at search data. A calendar reminder for the day after launch, and another a week later, closes that gap for the cost of two entries.

One last practical detail: the old site's redirects were reviewed too, because a launch that changes addresses and blocks crawlers at the same time doubles the damage. With the block lifted and the redirects checked, the agency's listings returned in roughly the order of how many links each page had earned over the years.

What would have caught it

PreviousThe CDN that served one customer's basket to anotherNextTwo nameservers, one provider, one bad afternoon

More from Hosting Autopsy

Autopsy

The wildcard certificate renewal that broke the mail server

An organisation used one wildcard certificate for its website and also installed the same files on its mail...

Autopsy

The mailbox that filled up and bounced the best client

A consultant used a single mailbox on her hosting plan for several years and never deleted anything....

Autopsy

The one-line .htaccess edit that took three sites down

An administrator added a redirect rule to a shared .htaccess file in a parent folder, which three sites...