robots.txt vs noindex vs canonical: Which to Use
Three tools, three different jobs. Use the wrong one and Google either keeps the page you wanted gone or never reads your instruction at all.
Use robots.txt to control crawling, noindex to keep a page out of search results, and rel=canonical to pick one URL from a set of duplicates. They solve different problems, and they get in each other's way: a page blocked in robots.txt can still be indexed, and Google never sees a noindex rule on a page it cannot crawl.
What is the difference between robots.txt, noindex and canonical?
Each one acts at a different stage. robots.txt decides whether a crawler may fetch a URL. noindex decides whether a fetched page may appear in results. A canonical tells Google which of several near-identical pages should represent the group. Here is how Google's documentation describes them side by side.
| robots.txt Disallow | noindex | rel=canonical | |
|---|---|---|---|
| Controls | Crawling | Indexing | Which duplicate represents the set |
| Where it lives | The robots.txt file at the site root | A robots meta tag or an X-Robots-Tag HTTP header | A link element in the head, or a Link HTTP header |
| Keeps a URL out of Google? | No. A disallowed URL can still be indexed if other sites link to it | Yes, once Googlebot crawls the page and reads the rule | No. Duplicates are consolidated, not removed |
| Needs the page to be crawled? | No | Yes | Yes |
| How firm is it? | Up to each crawler to obey | Google drops the page entirely | A hint, not a rule |
When should you use robots.txt?
When the problem is crawl traffic. Google's robots.txt introduction says the file is used mainly to avoid overloading your site with requests, and for web pages it suits two cases: your server is struggling with Google's crawler, or you want to stop crawling of unimportant or similar pages. It is also the documented way to keep image, video and audio files out of Google Search.
What it is not, in Google's words, is a mechanism for keeping a web page out of Google. A disallowed page can still be indexed when other sites link to it. The URL and anchor text from those links can show up in results, just without a description. And the rules are voluntary: Googlebot obeys them, other crawlers might not. Check what your file actually allows with our robots.txt tester.
When should you use noindex?
When a page must not appear in Google Search. Put <meta name="robots" content="noindex"> in the head, or send X-Robots-Tag: noindex as a response header for files such as PDFs. Google's noindex documentation says that once Googlebot crawls the page and extracts the rule, Google drops it from results entirely, regardless of whether other sites link to it.
The catch is in that first step. Google has to crawl the page to see meta tags and HTTP headers. If the page still shows up after you add noindex, the docs name two likely causes: Google has not recrawled it yet, or robots.txt is blocking it. The URL Inspection tool in Search Console can request a recrawl. Google also notes that some other search engines might interpret noindex differently.
When should you use a canonical tag instead?
When the same content lives at several URLs and you want one of them to collect the credit. Google's page on consolidating duplicate URLs lists the reasons: choose which URL people see in results, consolidate signals such as links into one preferred URL, simplify tracking metrics, and avoid spending crawl time on duplicates.
A canonical does not hide anything. Google's canonicalization page calls it a hint, not a rule, and says the canonical page is crawled most regularly while duplicates are crawled less often. A result usually points to the canonical, unless a duplicate suits the searcher better. Redirects and rel=canonical are both strong signals, and sitemap inclusion is a weak one. For non-HTML files, send the canonical as a Link HTTP header. Our guide to canonical tag conflicts covers what happens when those signals disagree.
Can you use robots.txt and noindex together?
Not on the same URL, if the goal is removal. Google is explicit: if a page is blocked by robots.txt, the crawler will never see the noindex rule, and the page can still appear in results. So the fix for a stubborn indexed URL is to remove the Disallow, let Google crawl it, and let the noindex do its work.
Google's duplicate-URL page names three more mismatches to avoid:
- Do not use robots.txt for canonicalization.
- Do not use noindex to steer canonical selection within one site, because it completely blocks the page from Search.
- Do not use the URL removal tool for canonicalization, because it hides all versions of a URL.
Which one should you reach for?
Start from the outcome you want, not the tool you know.
- Your server is overloaded, or crawlers waste time on unimportant pages: robots.txt.
- A page must not appear in Google: noindex, with crawling left open. For anything private, password protection, which Google names as the other option.
- The same content sits at several URLs: rel=canonical on each duplicate, or a redirect when the duplicate does not need to exist at all.
- Images, video or audio you want out of Google Search: robots.txt.
If AI crawlers are the concern, our robots.txt for AI crawlers guide has copy-paste rules per bot.