There are three kinds of AI bot: training crawlers that harvest content to build future models, retrieval bots that fetch a page live to answer a query and cite it, and agentic fetchers that only fire when a person pastes your URL. Allow the retrieval bots if you want to be cited, whatever you decide about the first category. Blocking the wrong one is the most common mistake in AI-search technical SEO.
Last verified: September 2026. Every row reflects the operator's own published documentation as of that date; no per-bot verification date is claimed, because none has been independently tested.
The registry table
| User agent | Operator | Category | Purpose | Honors robots.txt? |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | Crawls to improve future model training | Documented yes |
| OAI-SearchBot | OpenAI | Retrieval | Live fetch for ChatGPT Search citations | Documented yes |
| ChatGPT-User | OpenAI | Agentic / on-demand | Fetches a page only when a user asks ChatGPT to open it | Documented yes |
| ClaudeBot | Anthropic | Training | Crawls for model training data | Documented yes |
| Claude-SearchBot | Anthropic | Retrieval | Live fetch for Claude's web search tool — see Claude citation statistics | Documented yes |
| Claude-User | Anthropic | Agentic / on-demand | Fetches a page a user explicitly references in a Claude conversation | Documented yes |
| Google-Extended | Training (opt-out token) | Controls use of content for Gemini/AI training, separate from indexing | Documented yes | |
| Googlebot | Indexing (feeds Search, AIO, AI Mode) | Classic web crawl; also the retrieval backbone for AI Overviews and AI Mode — which do not cite the same sources | Documented yes | |
| PerplexityBot | Perplexity | Retrieval + training | Crawls and fetches for live answers and indexing | Disputed — see below |
| Perplexity-User | Perplexity | Agentic / on-demand | Fetches a page a user references directly | Documented yes |
| Bingbot | Microsoft | Indexing (feeds Copilot) | Classic crawl; backbone for Bing Copilot and (historically) ChatGPT Search | Documented yes |
| Mistral-User / MistralAI-User | Mistral | Agentic / retrieval | Fetches pages for Le Chat's web-aware features | Documented yes |
| Amazonbot | Amazon | Training + retrieval | Crawls for Alexa and Amazon AI features | Documented yes |
| Applebot / Applebot-Extended | Apple | Indexing + training (opt-out token) | Siri, Spotlight, and Apple Intelligence features | Documented yes |
| Meta-ExternalAgent | Meta | Training | Crawls for Meta AI model training | Documented yes |
| Diffbot | Diffbot | Retrieval (third-party, feeds multiple AI products) | Structured-extraction crawler used by several downstream AI tools | Documented yes |
"Documented yes" reflects the operator's own published policy. It is not independently verified compliance — see documented vs verified for what has actually been tested.
- The registry tracks 16 bots across 10 operators, verified against operator documentation in September 2026. Ten of the sixteen belong to just four companies, which is why one blanket rule cannot express a coherent policy toward any of them.
- Blocking a training crawler affects future models only. Blocking a retrieval bot removes you from that engine's citations on the next query. These are separate decisions with asymmetric costs.
- There is no universal "block all AI" token. Every bot has to be named individually in robots.txt.
- The compliance column reports published policy. Zero of the 16 has had its robots.txt compliance independently verified at scale, and Cloudflare has documented one operator using undeclared crawlers to evade no-crawl directives.
- This is not a complete list of AI crawlers. It covers the operators documented well enough to support a defensible per-bot recommendation; regional and enterprise-only crawlers are out of scope.
Training vs retrieval vs agentic
Nearly every mistake in this area comes from treating "AI bot" as one category. It is at least three, and the differences decide what a block actually costs you.
| Training crawler | Retrieval bot | Agentic fetcher | |
|---|---|---|---|
| Triggered by | The operator's own schedule | A user's query, at answer time | A user pasting or naming a URL |
| Typical volume | High, bulk, sustained | Moderate, query-driven | Low and spiky |
| Blocking it costs you | Influence on future models only | Citations in that engine, immediately | Almost nothing |
| Blocking it protects | Content from training reuse | Nothing you likely want protected | Very little |
| Reversible? | Forward-looking only — past crawls stand | Yes, on the next query | Yes |
| Example | GPTBot, ClaudeBot, Meta-ExternalAgent | OAI-SearchBot, Claude-SearchBot, PerplexityBot | ChatGPT-User, Claude-User, Perplexity-User |
Blocking a training crawler affects whether your content shapes a future model. It has no bearing on whether that company's product can cite you today. Blocking a retrieval bot removes you from that engine's citations on the next query it answers.
Copy-paste robots.txt recipes
Three common intents. Paste the block that matches yours into the robots.txt of each host you run, then confirm from your logs that the bot re-fetched the file. If an agent is writing the file for you, read the agentic SEO failure catalogue first — over-blocking robots.txt is the one failure it rates critical, because nothing in a passing build catches it.
1. Allow AI-search citation, block AI training — the configuration most sites that want visibility end up with.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: / 2. Block everything AI-related. There is no wildcard token that reliably covers every operator, so each is named. Note that this also forecloses citation in every engine listed.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
User-agent: Claude-User
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Perplexity-User
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: MistralAI-User
Disallow: /
User-agent: Diffbot
Disallow: / 3. Allow everything. No entry is needed — the absence of a rule means allow. This is the default if you take no action.
A wildcard Disallow: / under User-agent: * blocks everything, Googlebot and your
organic traffic included, which is rarely the actual intent. OpenAI documents its three tokens separately
in its own bots reference,
and the same is true of every other operator in the table.
Which policy fits your site
- 01 Identify the bot Match the user agent against this registry
- 02 Classify its purpose Training, retrieval, or agentic/on-demand
- 03 Decide per-bot Not a single blanket policy for "AI"
- 04 Write the robots.txt rule Named explicitly — no universal wildcard exists
- 05 Verify it worked Check logs for continued (or stopped) requests from that agent
The registry does not prescribe one right answer. A publisher that wants citation reach allows the three retrieval bots and blocks the training crawlers. A business prioritising licensing revenue might block everything and pursue direct commercial deals. A store weighing AI-assisted product discovery against catalog copy being reproduced elsewhere may allow retrieval and block training more aggressively — a calculation that depends on what AI-sourced visits are actually worth and on whether citation reach is realistically available in your category at all.
The two failure directions have asymmetric costs. Blocking a retrieval bot by mistake silently removes a site from that engine's citations, with no error and no notification. Allowing a training bot you meant to block costs nothing visible but forecloses a future option, because blocking afterward does not retroactively remove what was already crawled.
How to verify a crawler is real
A user-agent string is a header any script can set. Spoofed "GPTBot" traffic is common enough that verification matters before you trust a log line — and Cloudflare has documented an operator itself using undeclared crawlers to evade no-crawl directives.
*.openai.com or *.crawl.anthropic.com pattern, per each operator's published docs).Why a correct-looking rule does nothing
Most failures here are mechanical rather than strategic. The intent is right and the rule still has no effect.
Per host, not per site. A crawler reads the robots.txt of the exact host it is requesting. A rule on the root domain does not govern a documentation or shop subdomain.
Groups are delimited by blank lines. A stray blank line splits a group, so rules that appear to belong to a named bot can silently end up somewhere else.
Matching is on the token. The bot compares its own declared token, not the whole user-agent string. GPTBot does not match OAI-SearchBot even though both belong to OpenAI. This single fact explains most of the mistakes this registry exists to prevent — the pair has its own deep dive in GPTBot vs OAI-SearchBot.
The most specific group wins. If a named group and a wildcard group both exist, the named one applies and the wildcard is ignored entirely for that bot. Blocking broadly and then allowing retrieval bots by name is a valid pattern.
The file is cached. Changes are not instant, and no operator in this table publishes its cache duration. The lag between an edit and its effect is the same variable measured in the crawl-to-citation latency study.
Nothing in this chain reports back to the site owner. There is no confirmation, no error, and no dashboard. Server logs are the only feedback loop that exists.
What is documented vs what is verified
The table above states documented policy. Independent, testable evidence of actual behaviour is much thinner, and this is a genuine research gap rather than a settled question. The crawl-to-referral ratios below are the most quoted figures in this space and among the least reproducible: Seomator's crawl-to-refer breakdown and Search Engine Journal's report that Googlebot still tops AI crawler traffic both derive from the same Cloudflare aggregate, so neither is independent confirmation of the other.
Whether every listed bot actually honors robots.txt in practice, versus only in documentation, has not been independently tested at scale.
Whether retrieval bots render JavaScript before extracting content — this determines whether client-rendered pages are visible to them at all.
Both are the kind of question a controlled test site can answer cheaply. The JavaScript-rendering question has its own pre-registered experiment at do AI crawlers render JavaScript?. Until either runs, treat the compliance column as policy, not proof. Anything tested and not confirmed lands in the null results registry rather than being quietly dropped.
Pitfalls when reading crawler logs
Verification tells you whether a request is genuine. It does not stop you drawing the wrong conclusion from a set of genuine requests.
Counting spoofed traffic. An unverified log line is a claim, not an observation.
Comparing bots on raw volume. A bulk training crawler and an on-demand fetcher are not comparable on request count. A low number from an on-demand bot is normal.
Reading crawls as visibility. Crawl volume and citation correlate loosely at best. A page can be crawled heavily and cited nowhere, which is why answer sampling rather than log analysis is the only way to measure citation.
Ignoring status codes. A thousand requests that all returned errors is a very different fact from a thousand successful fetches.
Short windows. Crawl cadence is uneven, and a single week can look like a collapse or a surge purely from scheduling. Prefer months to weeks.
There is no public, standardized methodology for AI-crawler log analysis. Every published figure in this space, including the ratios charted above, rests on methodological choices the reader cannot inspect. Treat cross-source comparisons with corresponding caution.
What would change this page
- Independent compliance testing at scale. If a third party systematically tested whether each bot honors a fresh block, the "documented yes" column could become an observed column. That would be the single largest improvement possible here.
- Evidence that blocking a training crawler affects citation. The registry assumes the two are independent, on the strength of operator documentation. Evidence of coupling would overturn the core recommendation.
- Evidence on JavaScript rendering by retrieval bots. If retrieval bots do not render, client-side pages are invisible to them regardless of robots.txt, which would make rendering a bigger lever than blocking.
- A widely adopted wildcard token for AI crawlers. That would collapse the per-bot recipes above into a single directive.
- Agentic protocols reaching real adoption. The WebMCP and agent-readable web layer describes a mode of interaction that is neither a bulk crawl nor a live retrieval nor a user-triggered fetch. Broad adoption would add a fourth category rather than a row.
Last verified: September 2026
- Restructured so the registry table sits above the fold; cut the two worked scenarios and the comparison-with-other-lists section.
- Added full copy-paste robots.txt blocks for the three common intents.
- Added a page-level last-verified date. No per-bot verification dates are claimed, because none exist.
- Expanded to 16 bots across 10 operators; added Meta-ExternalAgent and Diffbot.
- Initial publication tracked 14 bots across 8 operators.
Limitations
- Operator documentation can lag actual behavior. A published policy is not a guarantee of current practice.
- New bots ship faster than any registry can track them. This list covers the operators large enough to matter for most sites, not every AI company crawling the web.
- Regional and enterprise-only bots are excluded — this registry is scoped to consumer-facing AI search and assistant products.
- Local and vertical coverage is out of scope — local AI search and YMYL categories may see different crawler behaviour this registry does not characterise.
- No compliance claim here is independently verified. The honest description of this page is a well-organized summary of published operator policy, with explicit flags wherever policy and verified behaviour have not been shown to match.
Next: check what your own site currently tells these bots. Paste your domain into the llms.txt generator to see the retrieval-bot directives you are publishing today, then reconcile them against the table above. If you would rather work through the whole access layer in order, the technical GEO audit starts there.
Namdev, R. (2026). The AI Bot User-Agent Registry (v5). Retrieved from https://ritiknamdev.com/blog/ai-bot-user-agent-registry Published under CC BY 4.0 — reuse freely with attribution.
This is the crawler-access pillar under AI SEO, and the layer everything else depends on: a bot that cannot fetch a page cannot cite it. See GPTBot vs OAI-SearchBot for the deep dive on the single most-confused pair in this table, and AI crawler statistics for the traffic economics behind the crawl-to-referral numbers below.