Get your free SEO audit today Call 91 060 30 90
Home / Blog / SEO & GEO
SEO & GEO

Robots.txt: the file that might be blocking your SEO without you knowing

The robots.txt file lives at the root of your domain (yourdomain.com/robots.txt) and gives instructions to search engine crawlers about which parts of the site they can crawl and which they can't. It's a plain text file, with no technical mystery to it, but with disproportionate power: one badly written directive can pull your entire website out of Google's crawl without any obvious warning, beyond watching your traffic slowly decline.

The most common and most painful case happens right after migrating a site or launching a redesign. During development, it's normal (and recommended) to block crawler access to the staging environment with a line like Disallow: /. The problem shows up when that same configuration gets uploaded by mistake to the real production domain, and you're suddenly telling Google not to visit anything on your site at all.

What robots.txt can and can't do

It can stop a crawler from accessing a specific folder or page, which is useful for private areas, admin panels, or duplicate content that adds no indexable value (like internal filter results on a store). What it doesn't do is remove a URL that's already indexed: if a page is already in Google and you later block it via robots.txt, it can keep showing up in results (sometimes with no description, just the URL), because blocking crawling isn't the same as deindexing. That's what the noindex tag is for, a different and complementary instruction.

The most common mistakes

Beyond the accidental full block, three mistakes come up a lot: accidentally blocking the folders where CSS and JavaScript files live (which stops Google from rendering the page correctly and understanding what it actually looks like), blocking URL parameters so broadly that it takes down pages you actually wanted indexed, and forgetting to declare the XML sitemap's location inside robots.txt itself, a reference that helps crawlers find it without depending on you only declaring it in Search Console.

How to check yours right now

Just type yourdomain.com/robots.txt into your browser and read the result, or use the robots.txt tester built into Google Search Console, which also tells you if a specific URL is blocked and by exactly which line in the file. It's a two-minute check worth doing after any migration, hosting change or redesign, precisely because that's when this kind of mistake tends to slip through.

Frequently asked questions

How do I know if my site has this problem right now?

Go into Google Search Console, under the crawling section, and check whether any pages are flagged as "blocked by robots.txt". If important pages show up on that list, you have the problem active right now.

Is blocking a page with robots.txt enough to hide it from Google?

Not necessarily, as explained above: the reliable way to stop a page from showing up in results is to use the noindex tag on the page itself, not to block its crawling via robots.txt.

Do I need a robots.txt even if I want Google to index my entire site?

Yes, it's still worth having one, even a very simple one, allowing full crawling and pointing to your sitemap. Its absence doesn't block anything, but having it properly configured is good practice that avoids ambiguity.

More on SEO & GEO

Shall we talk about seo & geo for your business?

Tell us about your project and we'll tell you how we can help, no strings attached.

Call 91 060 30 90