A regional estate agency commissioned a new website from a freelance developer. While building it on a temporary address, the developer did what careful developers do: she did not want half-finished pages appearing in search results. So she ticked the WordPress option that asks search engines not to index the site, and she also added a robots.txt file that disallowed everything.
When the site went live, the files were copied across with the rest. Nobody remembered either of them, because both were invisible. A page with a robots rule looks exactly like a page without one.
Two switches that look like one
The block came from two places, and that is part of why it survived. The first was the WordPress setting under Settings, Reading, labelled "Discourage search engines from indexing this site". When ticked, WordPress adds a noindex instruction to every page, in a meta tag in the HTML and in an HTTP header on some versions. The second was a hand-written file at the root of the site:
User-agent: *
Disallow: /
Those two lines mean: every crawler, stay out of everything. A well-behaved search engine reads the file before fetching anything, sees the rule, and does not request the pages.
The two mechanisms do different jobs, and the difference matters when you come to repair things. A robots rule controls crawling: whether the bot may fetch the page. A noindex instruction controls indexing: whether a page it has fetched may appear in results. They interact in an awkward way. If robots.txt forbids the fetch, the crawler never sees the page, so it never sees a noindex either, and a blocked address can in some cases still appear in results as a bare link with no description, based on links from elsewhere.
What the owner noticed, and when
For the first days nothing seemed wrong. Visitors who typed the address or followed the old bookmarks reached the new site, the phone rang as usual, and the agency's staff looked at the site many times a day without discovering anything. The old site's pages were already in the index, and the search engines kept showing them for a while, with the old addresses leading to redirects or errors.
A week after launch, the owner searched for the company name and saw only old results, and one with a description reading "No information is available for this page." Analytics showed organic visits falling off a cliff: roughly 400 a day became about 60, almost all of them returning visitors who knew the name. The new site had been live and invisible for nine days.
The phrase that gave the game away appeared when the owner clicked that odd search result: the search engine could not show a description because it was not allowed to read the page. That is the visible symptom of a robots block, and it is easy to miss unless you know it.
Finding it
The first guess was a penalty for changing the site, which is a popular theory and rarely right. The second was that the new site was too slow. The developer, called back, thought of the staging settings in about thirty seconds, which says something about the value of asking the person who built it.
The check itself takes seconds:
curl -s https://example.com/robots.txt
curl -sI https://example.com/ | grep -i x-robots-tag
curl -s https://example.com/ | grep -i "name=.robots."
The first command showed the two lines above. The third showed <meta name='robots' content='noindex, nofollow' /> in the page head. Two blocks, both left from the build.
The search console, which the agency had not opened since the old site's days, would have said it plainly under its indexing reports: pages blocked by robots.txt, excluded by noindex tag. It also offers a tester that fetches a page the way the crawler would.
Removing the block, and why recovery took weeks
Removing the block was quick. The settings box was unticked and the robots file was replaced by one that allowed everything and pointed to the sitemap:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/sitemap.xml
Getting back to where they had been took several weeks. Search engines do not re-check a robots file on every request; they cache it for a while, often up to a day, and then they need to recrawl the pages, which they schedule by how important they judge the site to be. Several important pages, such as the area guides that had brought in most of the traffic, had been dropped from the index during the nine days and had to be found again from scratch. Submitting the sitemap and requesting indexing of the key pages in the search console sped it up, but did not make it instant.
Commands worth running
Make this a launch routine and repeat it after any migration:
- Open
/robots.txtin a browser and read every line. If it saysDisallow: /underUser-agent: *, stop. - In WordPress, look at Settings, Reading. The search visibility box must be unticked on the live site.
- View the page source of the homepage and search for
noindex. - Run
curl -Ion a few pages and look for anX-Robots-Tagheader, which some servers and plugins add. - Submit the sitemap in the search console and open the indexing report a day or two later.
Also protect the staging site properly. A password on the whole staging address (HTTP authentication) keeps out both crawlers and curious visitors, and it is not copied to the live site by accident the way a file is. For the status codes you may meet in these checks, see the status code reference.
| Method | Stops crawling | Stops listing | Travels to live by mistake |
|---|---|---|---|
| robots.txt Disallow | Yes | Not reliably | Yes, it is a file |
| Noindex tag or setting | No | Yes | Yes, if in the database |
| HTTP password on staging | Yes | Yes | No, it belongs to the server |
Quick answers
Is robots.txt a security measure?
No. It is a polite request, and the file itself is public. Never list private paths in it as if that hid them.
How long until the pages return?
For a small site, days to a few weeks. Important pages with many links tend to come back first.
Can I have an empty robots.txt?
Yes, an empty file means everything is allowed, and so does having none. A sitemap line is a useful addition.
Why does a blocked page still appear in results?
Because other sites link to it, the search engine may list the address without reading it. Allow crawling and use noindex if you want it gone.
What the agency changed afterwards
The agency agreed a rule with its developer: staging sites live behind a password, never behind a robots file alone. The launch checklist, now a shared document, has a line for the search visibility box, a line for robots.txt and a line for opening the homepage source and searching for noindex. Each line has a name and a date next to it.
They also learnt something about the timing of monitoring. The nine days were not a failure of any one person; they were a gap between the moment the site went live and the moment anybody looked at search data. A calendar reminder for the day after launch, and another a week later, closes that gap for the cost of two entries.
One last practical detail: the old site's redirects were reviewed too, because a launch that changes addresses and blocks crawlers at the same time doubles the damage. With the block lifted and the redirects checked, the agency's listings returned in roughly the order of how many links each page had earned over the years.
What would have caught it
- A launch checklist with a line for robots.txt and the WordPress visibility setting, signed off by a person who did not build the site.
- Fetching the live
robots.txtand the homepage source yourself, within an hour of going live. - Protecting the staging site with a password at the server, so that no block files need to exist at all.
- A search console alert for sudden indexing drops, with the account owned by the business and not only the developer.
- A weekly glance at organic visits for the first month after any relaunch.
- A scheduled check that fetches
/robots.txtdaily and warns if it ever contains a blanket Disallow.