Overview
Three mechanisms act at three stages: robots.txt decides whether a crawler may fetch a URL, the robots meta tag and X-Robots-Tag header decide whether a fetched URL is indexed and served, and snippet directives decide how much of an indexed page may be shown. Using the first for the second job is the usual reason “blocked” pages stay in results. RFC 9309 defines robots.txt; Google and Bing document the rest.
Pick the mechanism by goal
Choose the mechanism that acts at the stage you want to change.
| Goal | Mechanism | Gotcha |
|---|---|---|
| Stop crawling a URL space | Disallow in robots.txt | A blocked URL that other sites link to can still be indexed. |
| Keep an HTML page out of results | <meta name="robots" content="noindex"> | The page must be crawlable; a Disallow hides the tag. |
| Keep a PDF, image, or video out | X-Robots-Tag: noindex response header | Needs server or CDN config; these files cannot carry a meta tag. |
| Index the page, show no snippet | nosnippet | At Google it also blocks use as direct input to AI Overviews and AI Mode. |
| Cap snippet or preview size | max-snippet, max-image-preview, max-video-preview | max-snippet:0 equals nosnippet; at -1 Google picks the length and Bing sets no limit. |
| Hide one passage from snippets | data-nosnippet | Do not add or remove it with JavaScript (Google). |
| Retire a page on a date | unavailable_after | Googlebot also crawls the URL far less after the date; return 404 or 410 once the page is gone. |
| Opt out of Gemini training and grounding | Google-Extended in robots.txt | A token, not a crawler; no effect on Search or AI Overviews. |
| Keep content private | Authentication | RFC 9309 says robots rules are not access authorization. |
Use robots.txt for crawl control only
robots.txt answers “may this crawler fetch this URL” and nothing else. Google calls it “not a mechanism for keeping a web page out of Google”.
- Serve one file per host, protocol, and port, at the root.
- Google reads only
user-agent,allow,disallow, andsitemap;noindexandcrawl-delaylines are ignored. - The longest matching path wins; on a conflict Google applies the least restrictive rule.
*and$are the only wildcards. - Google ignores everything past 500 KiB.
- A 4xx (except 429) counts as no restrictions. On a 5xx Google stops crawling for 12 hours, then uses the last good copy for up to 30 days. Google caches the file up to 24 hours.
User-agent: *
Disallow: /search
Disallow: /*?sort=
Allow: /search/help
Sitemap: https://example.com/sitemap.xml
Which URL spaces to block: crawl-budget and pagination-and-facets. Root-file catalog: discoverability-files.
Keep a noindex page crawlable
noindex works only when the crawler can fetch the page. Google: when robots.txt blocks the URL, “the crawler will never see the noindex rule, and the page can still appear in search results”. Bing’s webmaster guidelines likewise say to allow Bingbot to crawl noindex pages.
- Never put
Disallowandnoindexon one URL. - To remove an already blocked URL, delete the
Disallow, servenoindex, and request a recrawl in the URL Inspection tool. noindexpages are still fetched. UseDisallowfor crawl waste andnoindexfor index control.
Use X-Robots-Tag for files without a head
Use the meta tag on HTML and the header on everything else; both take the same rules and are read only when the URL is crawled, so a change applies at the next recrawl.
<meta name="robots" content="noindex, nosnippet">location ~* \.pdf$ {
add_header X-Robots-Tag "noindex";
}curl -sI https://example.com/report.pdf | grep -i x-robots-tag- Names and values are case-insensitive; combine rules with commas or separate tags. The more restrictive rule wins on conflict.
- Staging pattern: seo-migration-playbook.
Use only directives the engine documents
Google’s supported rules:
| Directive | Effect |
|---|---|
noindex | Keep the page, media, or resource out of results. |
nofollow | Do not follow the page’s links; Google may otherwise use them for discovery. |
none | Same as noindex, nofollow. |
nosnippet | No text snippet or video preview. |
max-snippet:N | Snippet cap in characters; 0 equals nosnippet. |
max-image-preview: | none, standard, or large. |
max-video-preview:N | Preview cap in seconds; 0 allows a static image only. |
indexifembedded | Only with noindex; lets Google index content embedded through an iframe. |
noimageindex | Do not index images on the page. |
notranslate | Do not offer translation in results. |
unavailable_after:date | Stop showing after the date; RFC 822, RFC 850, or ISO 8601. |
all | No restrictions. |
Bing documents noindex, nofollow, nosnippet, noarchive, nocache, max-snippet, max-image-preview, max-video-preview, and data-nosnippet, and its 2020 announcement says it treats the max-* options as directives, not hints. Treat the rest as Google-only. Google no longer uses noarchive.
Hide passages with data-nosnippet
Use data-nosnippet to drop one passage from snippets while the page stays indexed and rankable.
<div data-nosnippet>Reader comments and hourly prices.</div>- Google accepts it on
span,div, andsection, needs valid HTML with closed tags, and warns against toggling it with JavaScript. - Bing added support in October 2025, accepts any element, and excludes the content from snippets and AI summaries while keeping it rankable.
Target one crawler with user-agent rules
Name the crawler on each surface when engines need different rules.
- In robots.txt a crawler obeys only its most specific matching group; groups are not combined. A
Googlebotgroup ignores the*group, so repeat shared rules. - In HTML use
name="googlebot"(Google also acceptsgooglebot-news). In the header writeX-Robots-Tag: googlebot: nofollow. A rule with no user agent applies to all crawlers. - Bing honors
crawl-delay; Google ignores it.
User-agent: *
Disallow: /private/
User-agent: Googlebot
Disallow: /private/
Disallow: /experiments/
Limit AI features with Search directives
Limit Google’s AI Overviews and AI Mode with the snippet and index directives that govern Search, because AI features are part of Search. Google lists nosnippet, data-nosnippet, max-snippet, and noindex as the controls and says no new machine-readable or AI text files are needed to appear in them. ai-txt and agents-txt are advisory policy files, not crawl controls.
Google-Extended is a standalone product token with no user agent string of its own. It governs whether crawled content trains future Gemini models and grounds Gemini Apps and Vertex AI, and it does not affect Search inclusion or ranking.
User-agent: Google-Extended
Disallow: /
Bing’s webmaster guidelines describe noindex as the control for search, Copilot, and grounding results, and noarchive as keeping content out of Copilot answers. Tokens for OpenAI, Anthropic, Perplexity, and others: ai-crawlers. Sites following the llm-discoverability-standard keep those tokens allowed.