robots.txt Mistakes That Quietly Break Your SEO

· Jonathon Roberts · 3 min read

A robots.txt file open in a code editor showing allow and disallow rules

robots.txt is a small file that causes an outsized amount of damage. It sits at the root of your site, it is usually a few lines long, and a single typo in it can remove hundreds of thousands of pages from Google's reach. I have seen entire migrations undone by one misplaced forward slash.

Here are the mistakes I see most often, and how to check whether your own file is doing what you think it is.

Blocking a page and expecting noindex to work

This is the most common one, and the most misunderstood. If a URL is blocked in robots.txt, Google cannot crawl it. If Google cannot crawl it, it cannot see the noindex tag on the page. The URL can still appear in search results, just without a snippet, because Google indexes the address rather than the content.

So a page that is both disallowed and noindexed can stay in the index indefinitely. If you want a page removed from search, let Google crawl it and see the noindex. Blocking it does the opposite of what most people intend.

Robots.txt syntax displayed on a monitor with user agent and disallow rules

Blocking the resources Google needs to render

Older sites often disallow directories like /js/, /css/ or /wp-includes/ because someone decided crawlers did not need them. Google renders pages now, not just reads HTML. If your stylesheet or your JavaScript bundle is blocked, Googlebot sees a broken or empty version of the page.

On a JavaScript-heavy site this can mean your content, your internal links and your structured data are all invisible, while the site looks perfectly fine in your browser. Search Console will flag this under Page Experience and in the URL Inspection render, but plenty of sites carry the problem for years without noticing.

Leaving staging rules on production

The classic. A staging site gets Disallow: / to keep it out of search. The site gets migrated to production, the file comes along for the ride, and the new site tells every crawler on the internet to go away. Rankings fall off a cliff over the next week and nobody connects it to the launch.

Any pre-launch checklist should include fetching yourdomain.com/robots.txt on the live domain, on a mobile connection or a VPN if you suspect caching, and reading it with your own eyes.

Expecting robots.txt to hide URLs

A disallow rule is a request, not a lock. Other crawlers can ignore it. Google usually honours it for crawling but, as above, the URL itself can still be indexed. If a page is genuinely sensitive, robots.txt is the wrong tool entirely. Use authentication, or noindex with the page left crawlable.

Assuming every section applies to every bot

Rules in robots.txt are grouped by user agent. A block that starts with User-agent: bingbot only applies to Bingbot. If you write a rule intended for everyone and put it under a specific bot, or leave a User-agent: * group with nothing in it, the rule does nothing.

Groups matter in the other direction too. Once a bot matches a more specific group, it ignores the * group completely. If you write rules for Googlebot and forget to repeat a disallow that exists under *, Googlebot is now allowed into the area you meant to keep closed.

A text file of directives open in a code editor

How to check yours

There is a robots.txt report in Search Console under Settings, and the URL Inspection tool will tell you whether a specific URL is blocked. Both are worth using rather than reading the file and guessing.

Beyond that, do two things. First, fetch the file yourself and check it returns what you expect, including on any subdomains you forgot existed. Second, run a crawl with Screaming Frog twice: once respecting robots.txt and once ignoring it. Any meaningful difference between the two lists is worth investigating, because it means part of your site is only visible when the rules are bypassed.

The file is tiny. The blast radius is not. Five minutes checking it after every release is one of the cheapest bits of technical SEO you can do.

More articles

All articles

How AI is Transforming SEO in 2026

The rapid evolution of artificial intelligence (AI) is revolutionising industries worldwide, and search engine optimisation (SEO) is no exception. By 2026, AI…