robots.txt: five mistakes that hide a whole site
Disallow is not noindex, and blocking a page stops Google reading the noindex on it. The five robots.txt errors that remove a site from search.
· 5 min read
What the file does, precisely
robots.txt is a plain text file at the root of a domain that tells well-behaved crawlers which paths they may request. That is the entire scope of it. It controls fetching, not indexing, and the distinction between those two things is the source of most of the damage the file causes. A Disallow line says 'do not request this'. It does not say 'do not list this in results', and it has no power to remove anything from an index.
It is also worth being clear that the file is a request rather than a security control. Compliant crawlers honour it; a scraper that does not care will read it as a helpful map of the directories you would prefer nobody visited. So a path you genuinely need kept private needs authentication, not a Disallow line. The file's legitimate uses are narrow and real: keeping crawlers out of endless faceted URL spaces, out of internal search result pages, out of endpoints that cost you money to serve. Those are crawl-budget decisions. None of them is a way to keep a page out of search results.
Mistake one: using Disallow when you meant noindex
This is the most common and the most consequential. Someone wants a page gone from search results, so they add a Disallow line for it. The page continues to appear, sometimes for months, occasionally with no description underneath because the crawler is no longer permitted to fetch the text it would have summarised. The owner concludes the file is not working and adds more lines.
The mechanism is straightforward once stated. A URL can be indexed on the strength of links pointing at it, without the page ever being fetched. Disallowing the fetch removes the engine's ability to see the page's content but not its knowledge that the URL exists. To remove a page from results you need a robots meta tag saying noindex in the page's own head, or an equivalent HTTP header — an instruction that lives in the response, which means the crawler has to be allowed to request the page in order to receive it. The correct tool depends on the goal: Disallow to save crawling, noindex to remove from results, and authentication to actually restrict access.
Mistake two: disallowing a page that carries a noindex
This is the trap that follows directly from the first, and it catches people who did the right thing and then did one more thing. A page is marked noindex correctly, the owner wants to be thorough, and they also add a Disallow for it in robots.txt. The result is that the page stays in results indefinitely, because the crawler is now forbidden from fetching the page and therefore never sees the noindex tag sitting in its head.
The two directives work against each other and the blocking one wins, in the sense that it prevents the other from ever being read. The sequence that actually works is to allow crawling, serve the noindex, wait for the page to be recrawled and dropped, and only then — if you have a crawl-budget reason — add a Disallow. Most of the time that last step is unnecessary. Anyone who has both directives on one URL and is waiting for it to disappear is waiting for something that cannot happen, and the fix is to remove the Disallow line and leave the noindex alone.
Mistake three: blocking the CSS and JavaScript
Older robots.txt files often carry lines blocking directories that hold stylesheets, scripts and image assets, sometimes inherited from a template written when crawlers did not execute JavaScript and fetching assets was considered waste. Google renders pages now. When it renders a page whose CSS and JavaScript it is not allowed to fetch, it sees something closer to unstyled markup, possibly with the main content missing entirely if that content is populated by a script.
The consequences are not limited to a wrong impression of your layout. Assessments that depend on how the page presents itself — whether it works on a phone, whether content sits above the fold, whether text is legible — are being made against a rendering that does not match what your visitors see. Neither the site nor the crawler reports an error, which is what makes this one persist: the page looks fine to everyone at the business, and the version being evaluated is a version nobody at the business has ever looked at. The check is quick, since you can read your own robots.txt and see whether any asset directory appears in it.
Mistake four: a wildcard that reaches further than intended
Pattern matching in robots.txt supports an asterisk for any sequence of characters and a dollar sign to anchor the end of a URL, and both are easy to get wrong in a direction that is far more destructive than expected. Disallow: /print* is probably intended to block printer-friendly views and will also block anything beginning with those characters — a directory called /printing-services, for instance, which may be a page you sell from.
The more painful version involves a trailing slash. Disallow: /blog blocks /blog and everything beneath it, while Disallow: /blog/ blocks the contents but not the /blog path itself, and the two are visually near-identical in a file nobody rereads. The most destructive single line is Disallow: / with no path after it, which blocks the entire site and is a legitimate line for a staging environment. Any pattern is worth testing against a handful of real URLs from your own site before it ships, including the ones you did not have in mind when you wrote it — the whole risk of a wildcard is the URLs you were not thinking about.
Mistake five: the staging file that shipped to production
Staging sites are routinely protected with a two-line robots.txt containing a user-agent line and Disallow: /, which is correct — you do not want an unfinished copy competing with the real site. The mistake happens at launch, when the whole environment is promoted to production and that file goes with it. The new site is live, looks perfect, and is asking every compliant crawler not to fetch any of it.
What makes this the worst of the five is how long it survives. Nothing on the site appears broken. Every page loads for every human visitor. There is no error in any log, and the only symptom is an absence — traffic that never starts arriving, which is difficult to distinguish from a new site simply taking time. Businesses have gone months like this. The prevention is a single item on the launch checklist: after the deploy, open the live domain's own robots.txt in a browser and read it. It is one URL and it takes seconds, and it is worth repeating after any migration or platform change, because a redeploy can quietly restore a file somebody had already fixed.
Common questions
What is the difference between Disallow and noindex?
Disallow, in robots.txt, asks a crawler not to fetch a URL. Noindex, in the page's own meta tag or HTTP header, tells a search engine not to list it. One controls requests, the other controls results, and using the first when you meant the second leaves the page in search results with no description.
Will blocking a page in robots.txt get it out of Google?
No, and it can make removal harder. The URL may already be known from links and can stay listed without being fetched. If the page also carries a noindex tag, blocking the fetch prevents that tag from ever being read, so the page persists. Allow the crawl and serve the noindex instead.
Should I block my CSS and JavaScript to save crawl budget?
No. Rendering needs those files, and without them a page can be assessed as broken or empty while looking perfect to every human visitor. If crawl volume is a genuine problem, the places to look are faceted URL spaces, internal search pages and expensive endpoints, not the assets the page needs to display.
How do I check my own file safely?
Open your live domain followed by /robots.txt in a browser and read every line, then take a handful of real URLs you care about and check each pattern against them by hand. Pay particular attention to trailing slashes and any bare Disallow: / left over from a staging environment.
Related pages