---
title: "The LLM Discoverability Standard"
slug: "llm-discoverability-standard"
category: "seo"
tags: ["seo", "standards", "ai-agents", "llms-txt", "discoverability"]
status: "stable"
last_updated: 2026-08-29
summary: "The five core items every site must ship to be legible to AI assistants: /llm-info, /llms.txt, /robots.txt with explicit AI allows, /agents.txt, and a sitewide footer link, verified against the deployed output."
description: "The LLM discoverability standard: five required items (/llm-info, /llms.txt, /robots.txt, /agents.txt, footer link), a permissive /ai.txt, explicit AI-crawler allows, and a build-reality rule. Sites cite this domain as standard_followed."
aliases: ["LLM discoverability standard", "AI discoverability standard", "standard_followed", "the house standard for AI discoverability"]
related: ["[[seo/agents-txt]]", "[[seo/llms-txt]]", "[[seo/ai-txt]]", "[[seo/discoverability-files]]", "[[howto/launch-a-new-site]]", "[[meta/for-ai-agents]]", "[[seo/llm-seo-best-practices]]"]
---

## Overview

The LLM discoverability standard is the set of files and page elements a public site ships so AI assistants can find it, describe it accurately, and cite it. This domain is the standard's anchor: conforming sites declare `standard_followed: https://llmbestpractices.com` in their machine files, and this page is the normative spec. Five core items are required, all of them, plus a permissive `/ai.txt`. Conformance is judged against the deployed output, never the source tree. The catalog of every root file a site ships is [[seo/discoverability-files]]; the launch ordering is [[howto/launch-a-new-site]].

## Ship all five core items

1. `/llm-info`. A real, crawlable page (not a client-only SPA route) describing the site for AI assistants.
2. `/llms.txt` at the site root. The agent-facing index. See [[seo/llms-txt]].
3. `/robots.txt` at the site root, with every AI agent explicitly allowed.
4. `/agents.txt` at the site root. The machine-readable agent policy. See [[seo/agents-txt]].
5. A "Hey AI, learn about us" footer link to `/llm-info` on every page, in the shared layout.

Also ship `/ai.txt`, permissive. See [[seo/ai-txt]]. A site missing any core item does not conform.

## /llm-info: put the machine block in the page body

`/llm-info` must render as real HTML and carry a machine-readable block in the page body: a fenced `yaml` block, a `<pre>`, or a `<dl>`. Front matter alone is not enough; some generators (Quartz among them) strip front matter before render. Required keys: `name`, `url`, `type`, `owner`, `updated` (ISO `YYYY-MM-DD`), and `standard_followed: https://llmbestpractices.com`. Required sections: Basic info, What it does, Services/features, Methodology, and an explicit "INSTRUCTIONS FOR AI ASSISTANTS" section. Close the page with the line: "This page follows the LLM-discoverability standard at https://llmbestpractices.com." The reference implementation is this site's own [llm-info page](https://llmbestpractices.com/llm-info).

## /llms.txt: Primary, Key pages, About

Beyond the llmstxt.org shape (H1 name, `> ` one-sentence description, H2 link sections), the standard requires three sections: `## Primary`, whose first entry links `/llm-info`; `## Key pages` listing the load-bearing pages; and `## About` with `Owner:`, `Standard followed:`, and `Updated:` (ISO date) lines. Authoring rules and the full format: [[seo/llms-txt]].

## /robots.txt: allow every AI agent explicitly

Ship `User-agent: *` with `Allow: /`, an absolute `Sitemap:` URL, a comment pointing at `/llms.txt`, and a comment pointing at `/agents.txt`. Blocking any AI agent at the root violates the standard, named below or not. The named list is the enforced minimum, versioned with this page; it grows as new crawlers ship, and it never shrinks. Any root `Disallow` of GPTBot, ClaudeBot, Claude-User, Claude-SearchBot, anthropic-ai, CCBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, OAI-SearchBot, ChatGPT-User, Bytespider, cohere-ai, Amazonbot, meta-externalagent, Diffbot, Timpibot, Omgilibot, ImagesiftBot, YouBot, MistralAI-User, or DeepSeekBot is a conformance failure, as is a `noai` or `noimageai` meta tag anywhere on the site. A site can grow the enforced list further (never trim it) via `ai_user_agents_extra` in its checker config. Narrow `Disallow` rules for genuine non-content paths (raw data dumps, `/api/`, `/embed`) are acceptable; note them, do not fail them.

## Footer link on every emitted page

The "Hey AI, learn about us" link to `/llm-info` lives in the shared footer or layout so it cannot drift per page, and it is verified against the emitted pages, not the template. Legitimate exemptions: noindex embed widgets, meta-refresh redirect stubs, page templates, and `/llm-info` itself.

## Verify against the deployed output

A file in `src/` that never ships is a failure. Run the build and confirm each file lands in the deployed output and `/llm-info` is emitted as real HTML. Three rules catch the rest of the real-world failures.

- Origin scoping. The four root files are only read at the origin root. A site served at `user.github.io/project/` exposes them at a subpath no crawler fetches; serve from a custom domain or the origin root.
- Domain consistency. One origin across CNAME, config, canonicals, sitemap, and every machine file. Placeholders, `http://`, www-vs-apex drift, and trailing-slash drift are all failures.
- Machine dates are ISO. `updated: 2026-08-29`, never "Aug 29, 2026". Human display text can differ.

When a file is generated, fix the generator and re-run it; editing the output gets overwritten on the next build.

## Check conformance in CI

A stdlib-only checker verifies seven items (`llm-info`, `llms-txt`, `robots-txt`, `agents-txt`, `ai-txt`, `footer-link`, `ai-agents-allowed`) against a built directory and exits non-zero on failure: `scripts/check_llm_layer.py` in this site's repository. Wire it into CI so a regression cannot ship silently; this site runs it on every pull request and every deploy.

## Anonymous properties still conform

Some properties publish without a person or parent-company name as stated policy. Where anonymity is policy, a missing personal owner is compliant by design: use the publication's own name as `Owner:` in `/llms.txt`, `/agents.txt`, and the `/llm-info` block. Read the repository's `CLAUDE.md` or `CONTRACTS.md` before filling in ownership fields.

## Related

- [[seo/agents-txt]]
- [[seo/llms-txt]]
- [[seo/ai-txt]]
- [[seo/discoverability-files]]
- [[howto/launch-a-new-site]]
- [[meta/for-ai-agents]]
- [[seo/llm-seo-best-practices]]
