Overview

Three mechanisms act at three stages: robots.txt decides whether a crawler may fetch a URL, the robots meta tag and X-Robots-Tag header decide whether a fetched URL is indexed and served, and snippet directives decide how much of an indexed page may be shown. Using the first for the second job is the usual reason “blocked” pages stay in results. RFC 9309 defines robots.txt; Google and Bing document the rest.

Pick the mechanism by goal

Choose the mechanism that acts at the stage you want to change.

GoalMechanismGotcha
Stop crawling a URL spaceDisallow in robots.txtA blocked URL that other sites link to can still be indexed.
Keep an HTML page out of results<meta name="robots" content="noindex">The page must be crawlable; a Disallow hides the tag.
Keep a PDF, image, or video outX-Robots-Tag: noindex response headerNeeds server or CDN config; these files cannot carry a meta tag.
Index the page, show no snippetnosnippetAt Google it also blocks use as direct input to AI Overviews and AI Mode.
Cap snippet or preview sizemax-snippet, max-image-preview, max-video-previewmax-snippet:0 equals nosnippet; at -1 Google picks the length and Bing sets no limit.
Hide one passage from snippetsdata-nosnippetDo not add or remove it with JavaScript (Google).
Retire a page on a dateunavailable_afterGooglebot also crawls the URL far less after the date; return 404 or 410 once the page is gone.
Opt out of Gemini training and groundingGoogle-Extended in robots.txtA token, not a crawler; no effect on Search or AI Overviews.
Keep content privateAuthenticationRFC 9309 says robots rules are not access authorization.

Use robots.txt for crawl control only

robots.txt answers “may this crawler fetch this URL” and nothing else. Google calls it “not a mechanism for keeping a web page out of Google”.

  • Serve one file per host, protocol, and port, at the root.
  • Google reads only user-agent, allow, disallow, and sitemap; noindex and crawl-delay lines are ignored.
  • The longest matching path wins; on a conflict Google applies the least restrictive rule. * and $ are the only wildcards.
  • Google ignores everything past 500 KiB.
  • A 4xx (except 429) counts as no restrictions. On a 5xx Google stops crawling for 12 hours, then uses the last good copy for up to 30 days. Google caches the file up to 24 hours.
User-agent: *
Disallow: /search
Disallow: /*?sort=
Allow: /search/help

Sitemap: https://example.com/sitemap.xml

Which URL spaces to block: crawl-budget and pagination-and-facets. Root-file catalog: discoverability-files.

Keep a noindex page crawlable

noindex works only when the crawler can fetch the page. Google: when robots.txt blocks the URL, “the crawler will never see the noindex rule, and the page can still appear in search results”. Bing’s webmaster guidelines likewise say to allow Bingbot to crawl noindex pages.

  • Never put Disallow and noindex on one URL.
  • To remove an already blocked URL, delete the Disallow, serve noindex, and request a recrawl in the URL Inspection tool.
  • noindex pages are still fetched. Use Disallow for crawl waste and noindex for index control.

Use X-Robots-Tag for files without a head

Use the meta tag on HTML and the header on everything else; both take the same rules and are read only when the URL is crawled, so a change applies at the next recrawl.

<meta name="robots" content="noindex, nosnippet">
location ~* \.pdf$ {
  add_header X-Robots-Tag "noindex";
}
curl -sI https://example.com/report.pdf | grep -i x-robots-tag
  • Names and values are case-insensitive; combine rules with commas or separate tags. The more restrictive rule wins on conflict.
  • Staging pattern: seo-migration-playbook.

Use only directives the engine documents

Google’s supported rules:

DirectiveEffect
noindexKeep the page, media, or resource out of results.
nofollowDo not follow the page’s links; Google may otherwise use them for discovery.
noneSame as noindex, nofollow.
nosnippetNo text snippet or video preview.
max-snippet:NSnippet cap in characters; 0 equals nosnippet.
max-image-preview:none, standard, or large.
max-video-preview:NPreview cap in seconds; 0 allows a static image only.
indexifembeddedOnly with noindex; lets Google index content embedded through an iframe.
noimageindexDo not index images on the page.
notranslateDo not offer translation in results.
unavailable_after:dateStop showing after the date; RFC 822, RFC 850, or ISO 8601.
allNo restrictions.

Bing documents noindex, nofollow, nosnippet, noarchive, nocache, max-snippet, max-image-preview, max-video-preview, and data-nosnippet, and its 2020 announcement says it treats the max-* options as directives, not hints. Treat the rest as Google-only. Google no longer uses noarchive.

Hide passages with data-nosnippet

Use data-nosnippet to drop one passage from snippets while the page stays indexed and rankable.

<div data-nosnippet>Reader comments and hourly prices.</div>
  • Google accepts it on span, div, and section, needs valid HTML with closed tags, and warns against toggling it with JavaScript.
  • Bing added support in October 2025, accepts any element, and excludes the content from snippets and AI summaries while keeping it rankable.

Target one crawler with user-agent rules

Name the crawler on each surface when engines need different rules.

  • In robots.txt a crawler obeys only its most specific matching group; groups are not combined. A Googlebot group ignores the * group, so repeat shared rules.
  • In HTML use name="googlebot" (Google also accepts googlebot-news). In the header write X-Robots-Tag: googlebot: nofollow. A rule with no user agent applies to all crawlers.
  • Bing honors crawl-delay; Google ignores it.
User-agent: *
Disallow: /private/

User-agent: Googlebot
Disallow: /private/
Disallow: /experiments/

Limit AI features with Search directives

Limit Google’s AI Overviews and AI Mode with the snippet and index directives that govern Search, because AI features are part of Search. Google lists nosnippet, data-nosnippet, max-snippet, and noindex as the controls and says no new machine-readable or AI text files are needed to appear in them. ai-txt and agents-txt are advisory policy files, not crawl controls.

Google-Extended is a standalone product token with no user agent string of its own. It governs whether crawled content trains future Gemini models and grounds Gemini Apps and Vertex AI, and it does not affect Search inclusion or ranking.

User-agent: Google-Extended
Disallow: /

Bing’s webmaster guidelines describe noindex as the control for search, Copilot, and grounding results, and noarchive as keeping content out of Copilot answers. Tokens for OpenAI, Anthropic, Perplexity, and others: ai-crawlers. Sites following the llm-discoverability-standard keep those tokens allowed.