Crawling and indexing · Practical guide

robots.txt vs noindex: which one should you use?

Crawling and indexing are separate decisions. robots.txt gives crawlers access instructions, while noindex tells a supporting search engine not to include a resource in its index. Choosing the right mechanism starts with stating exactly what you want to happen to the URL.

Use robots.txt to control crawling

robots.txt contains crawling rules for user agents. A disallowed URL can still appear in Google’s results if it is discovered through other links, even though Google cannot fetch its content. A crawl block is therefore not a reliable instruction to remove a page from search.

Robots rules are public and are not access control. If content must be private, protect access to it. Do not use a list of sensitive paths in robots.txt as a substitute for authentication.

Reference: Google Search Central: Introduction to robots.txt

Use noindex when the page should stay out of search

For Google to apply noindex, it must be allowed to crawl the resource and read the instruction. An HTML page can use a robots meta tag; an HTTP response can use X-Robots-Tag. Google does not support putting noindex in robots.txt.

If you block a page in robots.txt and place noindex in its HTML, the crawler may never see the noindex instruction. For removal, consider the complete access and indexing setup rather than adding both controls and assuming they reinforce each other.

<meta name="robots" content="noindex">

Reference: Google Search Central: Block indexing with noindex

Inspect the actual directives

Review the page’s reported robots field alongside its response headers and the robots.txt evidence. A deployment template may accidentally carry a staging restriction onto public pages. Conversely, an intentionally excluded utility page may be working exactly as intended.

SiteAudit retrieves evidence but does not simulate every crawler’s rule evaluation or confirm Google’s indexed state. It also does not render scripts that might change the page. Use Search Console to investigate how Google sees a URL you own.

Choose the control that matches the goal

Write the desired outcome in plain language first. “Users must sign in,” “crawlers should avoid this path” and “this public page should not appear in search” are different requirements.

  1. For private content, implement access control and verify it without relying on crawler cooperation.
  2. For crawl management, review robots.txt rules for the intended user agent and path.
  3. For a public page excluded from indexing, make noindex readable to the crawler.
  4. Recheck the deployed response and use Search Console for indexing follow-up.

Keep reading

Related guides