Reference · Maintained registry

The AI Bot User-Agent Registry

Every AI crawler that matters for a website owner in one place — what it's for, whose it is, how to verify it, and what to put in robots.txt. Maintained, not a one-off listicle.

Most 'AI crawler list' articles are static and go stale within months. This one is a living registry: dated, versioned, and updated whenever an operator changes behavior — with a changelog at the bottom rather than silent edits.

Ritik Namdev Ritik Namdev ·Published September 2026 ·16 bots tracked ·12 min read ·Last verified September 2026
The short version

There are three kinds of AI bot: training crawlers that harvest content to build future models, retrieval bots that fetch a page live to answer a query and cite it, and agentic fetchers that only fire when a person pastes your URL. Allow the retrieval bots if you want to be cited, whatever you decide about the first category. Blocking the wrong one is the most common mistake in AI-search technical SEO.

Last verified: September 2026. Every row reflects the operator's own published documentation as of that date; no per-bot verification date is claimed, because none has been independently tested.

The registry table

16 AI bots relevant to website owners, across 10 operators. Last verified September 2026.
User agentOperatorCategoryPurposeHonors robots.txt?
GPTBotOpenAITrainingCrawls to improve future model trainingDocumented yes
OAI-SearchBotOpenAIRetrievalLive fetch for ChatGPT Search citationsDocumented yes
ChatGPT-UserOpenAIAgentic / on-demandFetches a page only when a user asks ChatGPT to open itDocumented yes
ClaudeBotAnthropicTrainingCrawls for model training dataDocumented yes
Claude-SearchBotAnthropicRetrievalLive fetch for Claude's web search tool — see Claude citation statisticsDocumented yes
Claude-UserAnthropicAgentic / on-demandFetches a page a user explicitly references in a Claude conversationDocumented yes
Google-ExtendedGoogleTraining (opt-out token)Controls use of content for Gemini/AI training, separate from indexingDocumented yes
GooglebotGoogleIndexing (feeds Search, AIO, AI Mode)Classic web crawl; also the retrieval backbone for AI Overviews and AI Mode — which do not cite the same sourcesDocumented yes
PerplexityBotPerplexityRetrieval + trainingCrawls and fetches for live answers and indexingDisputed — see below
Perplexity-UserPerplexityAgentic / on-demandFetches a page a user references directlyDocumented yes
BingbotMicrosoftIndexing (feeds Copilot)Classic crawl; backbone for Bing Copilot and (historically) ChatGPT SearchDocumented yes
Mistral-User / MistralAI-UserMistralAgentic / retrievalFetches pages for Le Chat's web-aware featuresDocumented yes
AmazonbotAmazonTraining + retrievalCrawls for Alexa and Amazon AI featuresDocumented yes
Applebot / Applebot-ExtendedAppleIndexing + training (opt-out token)Siri, Spotlight, and Apple Intelligence featuresDocumented yes
Meta-ExternalAgentMetaTrainingCrawls for Meta AI model trainingDocumented yes
DiffbotDiffbotRetrieval (third-party, feeds multiple AI products)Structured-extraction crawler used by several downstream AI toolsDocumented yes

"Documented yes" reflects the operator's own published policy. It is not independently verified compliance — see documented vs verified for what has actually been tested.

What this page establishes
  • The registry tracks 16 bots across 10 operators, verified against operator documentation in September 2026. Ten of the sixteen belong to just four companies, which is why one blanket rule cannot express a coherent policy toward any of them.
  • Blocking a training crawler affects future models only. Blocking a retrieval bot removes you from that engine's citations on the next query. These are separate decisions with asymmetric costs.
  • There is no universal "block all AI" token. Every bot has to be named individually in robots.txt.
  • The compliance column reports published policy. Zero of the 16 has had its robots.txt compliance independently verified at scale, and Cloudflare has documented one operator using undeclared crawlers to evade no-crawl directives.
  • This is not a complete list of AI crawlers. It covers the operators documented well enough to support a defensible per-bot recommendation; regional and enterprise-only crawlers are out of scope.

Training vs retrieval vs agentic

Nearly every mistake in this area comes from treating "AI bot" as one category. It is at least three, and the differences decide what a block actually costs you.

Training crawlerRetrieval botAgentic fetcher
Triggered byThe operator's own scheduleA user's query, at answer timeA user pasting or naming a URL
Typical volumeHigh, bulk, sustainedModerate, query-drivenLow and spiky
Blocking it costs youInfluence on future models onlyCitations in that engine, immediatelyAlmost nothing
Blocking it protectsContent from training reuseNothing you likely want protectedVery little
Reversible?Forward-looking only — past crawls standYes, on the next queryYes
ExampleGPTBot, ClaudeBot, Meta-ExternalAgentOAI-SearchBot, Claude-SearchBot, PerplexityBotChatGPT-User, Claude-User, Perplexity-User
Fact

Blocking a training crawler affects whether your content shapes a future model. It has no bearing on whether that company's product can cite you today. Blocking a retrieval bot removes you from that engine's citations on the next query it answers.

Copy-paste robots.txt recipes

Three common intents. Paste the block that matches yours into the robots.txt of each host you run, then confirm from your logs that the bot re-fetched the file. If an agent is writing the file for you, read the agentic SEO failure catalogue first — over-blocking robots.txt is the one failure it rates critical, because nothing in a passing build catches it.

1. Allow AI-search citation, block AI training — the configuration most sites that want visibility end up with.

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

2. Block everything AI-related. There is no wildcard token that reliably covers every operator, so each is named. Note that this also forecloses citation in every engine listed.

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: MistralAI-User
Disallow: /

User-agent: Diffbot
Disallow: /

3. Allow everything. No entry is needed — the absence of a rule means allow. This is the default if you take no action.

A wildcard Disallow: / under User-agent: * blocks everything, Googlebot and your organic traffic included, which is rarely the actual intent. OpenAI documents its three tokens separately in its own bots reference, and the same is true of every other operator in the table.

Which policy fits your site

Want to be cited by an engine?Allow its retrieval bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot).
Don't want your content training future models?Block the training crawler (GPTBot, ClaudeBot) — this does not affect citation.
Fine with agentic fetchers?ChatGPT-User, Claude-User and similar only fire when a user pastes your URL — usually safe to allow.
Unsure about a new or unlisted bot?Default to allow for retrieval-labeled bots, block for anything unclear that crawls in bulk.
How to apply the registry
  1. 01 Identify the bot Match the user agent against this registry
  2. 02 Classify its purpose Training, retrieval, or agentic/on-demand
  3. 03 Decide per-bot Not a single blanket policy for "AI"
  4. 04 Write the robots.txt rule Named explicitly — no universal wildcard exists
  5. 05 Verify it worked Check logs for continued (or stopped) requests from that agent

The registry does not prescribe one right answer. A publisher that wants citation reach allows the three retrieval bots and blocks the training crawlers. A business prioritising licensing revenue might block everything and pursue direct commercial deals. A store weighing AI-assisted product discovery against catalog copy being reproduced elsewhere may allow retrieval and block training more aggressively — a calculation that depends on what AI-sourced visits are actually worth and on whether citation reach is realistically available in your category at all.

The two failure directions have asymmetric costs. Blocking a retrieval bot by mistake silently removes a site from that engine's citations, with no error and no notification. Allowing a training bot you meant to block costs nothing visible but forecloses a future option, because blocking afterward does not retroactively remove what was already crawled.

How to verify a crawler is real

A user-agent string is a header any script can set. Spoofed "GPTBot" traffic is common enough that verification matters before you trust a log line — and Cloudflare has documented an operator itself using undeclared crawlers to evade no-crawl directives.

Step 1Reverse-DNS the request IP. Legitimate bots resolve to an operator-owned hostname (an *.openai.com or *.crawl.anthropic.com pattern, per each operator's published docs).
Step 2Forward-DNS that hostname back to an IP and confirm it matches the original request IP. This double-lookup is the standard defense against DNS spoofing.
Step 3Cross-check against the operator's published IP ranges where available, rather than trusting the hostname pattern alone.
Automate itCloudflare and most CDNs perform this verification automatically and expose a "verified bot" flag in logs — the fastest path if you are already behind one. Scripted versions of the checks are covered in the Claude Code for SEO guide.

Why a correct-looking rule does nothing

Most failures here are mechanical rather than strategic. The intent is right and the rule still has no effect.

Per host, not per site. A crawler reads the robots.txt of the exact host it is requesting. A rule on the root domain does not govern a documentation or shop subdomain.

Groups are delimited by blank lines. A stray blank line splits a group, so rules that appear to belong to a named bot can silently end up somewhere else.

Matching is on the token. The bot compares its own declared token, not the whole user-agent string. GPTBot does not match OAI-SearchBot even though both belong to OpenAI. This single fact explains most of the mistakes this registry exists to prevent — the pair has its own deep dive in GPTBot vs OAI-SearchBot.

The most specific group wins. If a named group and a wildcard group both exist, the named one applies and the wildcard is ignored entirely for that bot. Blocking broadly and then allowing retrieval bots by name is a valid pattern.

The file is cached. Changes are not instant, and no operator in this table publishes its cache duration. The lag between an edit and its effect is the same variable measured in the crawl-to-citation latency study.

Fact

Nothing in this chain reports back to the site owner. There is no confirmation, no error, and no dashboard. Server logs are the only feedback loop that exists.

What is documented vs what is verified

The table above states documented policy. Independent, testable evidence of actual behaviour is much thinner, and this is a genuine research gap rather than a settled question. The crawl-to-referral ratios below are the most quoted figures in this space and among the least reproducible: Seomator's crawl-to-refer breakdown and Search Engine Journal's report that Googlebot still tops AI crawler traffic both derive from the same Cloudflare aggregate, so neither is independent confirmation of the other.

Unverified

Whether every listed bot actually honors robots.txt in practice, versus only in documentation, has not been independently tested at scale.

Open question
Unverified

Whether retrieval bots render JavaScript before extracting content — this determines whether client-rendered pages are visible to them at all.

Open question

Both are the kind of question a controlled test site can answer cheaply. The JavaScript-rendering question has its own pre-registered experiment at do AI crawlers render JavaScript?. Until either runs, treat the compliance column as policy, not proof. Anything tested and not confirmed lands in the null results registry rather than being quietly dropped.

Pitfalls when reading crawler logs

Verification tells you whether a request is genuine. It does not stop you drawing the wrong conclusion from a set of genuine requests.

Counting spoofed traffic. An unverified log line is a claim, not an observation.

Comparing bots on raw volume. A bulk training crawler and an on-demand fetcher are not comparable on request count. A low number from an on-demand bot is normal.

Reading crawls as visibility. Crawl volume and citation correlate loosely at best. A page can be crawled heavily and cited nowhere, which is why answer sampling rather than log analysis is the only way to measure citation.

Ignoring status codes. A thousand requests that all returned errors is a very different fact from a thousand successful fetches.

Short windows. Crawl cadence is uneven, and a single week can look like a collapse or a surge purely from scheduling. Prefer months to weeks.

Open question

There is no public, standardized methodology for AI-crawler log analysis. Every published figure in this space, including the ratios charted above, rests on methodological choices the reader cannot inspect. Treat cross-source comparisons with corresponding caution.

What would change this page

  • Independent compliance testing at scale. If a third party systematically tested whether each bot honors a fresh block, the "documented yes" column could become an observed column. That would be the single largest improvement possible here.
  • Evidence that blocking a training crawler affects citation. The registry assumes the two are independent, on the strength of operator documentation. Evidence of coupling would overturn the core recommendation.
  • Evidence on JavaScript rendering by retrieval bots. If retrieval bots do not render, client-side pages are invisible to them regardless of robots.txt, which would make rendering a bigger lever than blocking.
  • A widely adopted wildcard token for AI crawlers. That would collapse the per-bot recipes above into a single directive.
  • Agentic protocols reaching real adoption. The WebMCP and agent-readable web layer describes a mode of interaction that is neither a bulk crawl nor a live retrieval nor a user-triggered fetch. Broad adoption would add a fourth category rather than a row.

Last verified: September 2026

September 2026
  • Restructured so the registry table sits above the fold; cut the two worked scenarios and the comparison-with-other-lists section.
  • Added full copy-paste robots.txt blocks for the three common intents.
  • Added a page-level last-verified date. No per-bot verification dates are claimed, because none exist.
September 2026 (earlier)
  • Expanded to 16 bots across 10 operators; added Meta-ExternalAgent and Diffbot.
  • Initial publication tracked 14 bots across 8 operators.

Limitations

  • Operator documentation can lag actual behavior. A published policy is not a guarantee of current practice.
  • New bots ship faster than any registry can track them. This list covers the operators large enough to matter for most sites, not every AI company crawling the web.
  • Regional and enterprise-only bots are excluded — this registry is scoped to consumer-facing AI search and assistant products.
  • Local and vertical coverage is out of scope — local AI search and YMYL categories may see different crawler behaviour this registry does not characterise.
  • No compliance claim here is independently verified. The honest description of this page is a well-organized summary of published operator policy, with explicit flags wherever policy and verified behaviour have not been shown to match.

Next: check what your own site currently tells these bots. Paste your domain into the llms.txt generator to see the retrieval-bot directives you are publishing today, then reconcile them against the table above. If you would rather work through the whole access layer in order, the technical GEO audit starts there.

How to cite this
Namdev, R. (2026). The AI Bot User-Agent Registry (v5). Retrieved from https://ritiknamdev.com/blog/ai-bot-user-agent-registry

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

This is the crawler-access pillar under AI SEO, and the layer everything else depends on: a bot that cannot fetch a page cannot cite it. See GPTBot vs OAI-SearchBot for the deep dive on the single most-confused pair in this table, and AI crawler statistics for the traffic economics behind the crawl-to-referral numbers below.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Cloudflare Radar — verified bots and AI crawlersradar.cloudflare.com OpenAI — GPTBot and OAI-SearchBot documentationplatform.openai.com/docs/bots OpenAI developer docs — bots referencedevelopers.openai.com/api/docs/bots Anthropic — crawler and agent documentationdocs.anthropic.com Anthropic Support — Does Anthropic crawl the web, and how do I block it?support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Anthropic — Web search tool documentationplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Google Search Central — Googlebot and Google-Extendeddevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers Google Search Central — Managing crawl budget for large sitesdevelopers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget AmICited — Google-Extended: what it does and whether to block itwww.amicited.com/blog/google-extended-what-it-does-should-you-block-it Perplexity — PerplexityBot documentationdocs.perplexity.ai Perplexity — crawler referencedocs.perplexity.ai/docs/resources/perplexity-crawlers Cloudflare — Perplexity is using stealth, undeclared crawlers to evade no-crawl directivesblog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives Cloudflare — From Googlebot to GPTBot: who is crawling your site in 2025blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Cloudflare — Crawlers, clicks and AI bots (training vs retrieval traffic)blog.cloudflare.com/crawlers-click-ai-bots-training Cloudflare Radar — 2025 Year in Reviewblog.cloudflare.com/radar-2025-year-in-review InfoQ — Cloudflare 2025 AI bot findingsinfoq.com/news/2025/12/cloudflare-2025-ai-bots Search Engine Journal — Cloudflare report: Googlebot tops AI crawler trafficwww.searchenginejournal.com/cloudflare-report-googlebot-tops-ai-crawler-traffic/563303 Seomator — Crawl-to-refer ratios for AI crawlers and LLM botsseomator.com/blog/crawl-to-refer-ratio-ai-crawlers-llm-bots Paul Calvano — AI bots and robots.txtpaulcalvano.com/2025-08-21-ai-bots-and-robots-txt Technology Checker — robots.txt AI crawler blocking reporttechnologychecker.io/blog/robots-txt-ai-crawlers-blocking-report BuzzStream — Publishers blocking AI crawlers studywww.buzzstream.com/blog/publishers-block-ai-study Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots Digital Applied — 30-day agentic crawler behaviour log studywww.digitalapplied.com/blog/agentic-crawler-behavior-30-day-site-log-study AgentLux — Agentic traffic is here: preparing for AI browsers and shopping agentsagentlux.ai/blog/agentic-traffic-is-here-how-websites-should-prepare-for-ai-browsers-and-shopping-agents Previsible — Agentic shoppingprevisible.io/seo-ai-news/agentic-shopping dev.to — The state of agentic AI standards in 2026 (MCP, A2A, WebMCP)dev.to/alexmercedcoder/the-state-of-agentic-ai-standards-in-2026-mcp-a2a-webmcp-osi-and-the-protocol-stack-taking-3o2l Bing Webmaster Blog — bingbot series: submitting URLs for fast indexingblogs.bing.com/webmaster/january-2019/bingbot-Series-Get-your-content-indexed-fast-by-now-submitting-up-to-10,000-URLs-per-day-to-Bing
FAQ

Frequently asked questions

Should I block every AI crawler?
Blocking a training crawler (GPTBot, or ClaudeBot in its training role) stops your content being used to train future models. It has no effect on whether that engine can cite you today. Blocking a retrieval bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot) removes you from that engine's citations immediately. Decide per-bot, not as a blanket policy.
How do I actually verify a crawler is who it claims to be?
User-agent strings can be spoofed by anyone. The reliable method is a reverse-DNS lookup on the request IP, followed by a forward lookup to confirm it resolves back. Each major operator publishes the hostname pattern to check against.
Does blocking a bot in robots.txt actually stop it?
Only for crawlers that honour robots.txt, and compliance is self-reported, not enforced. The "documented vs verified" section states plainly what has been observed rather than merely published.
Why does a rule I added yesterday still seem to be ignored?
Operators cache robots.txt, and no operator in this table publishes how long its cache lasts. Before concluding a bot is non-compliant, confirm from your logs that it has re-fetched the file since your edit.
If I block a training crawler today, can I un-block it later without losing anything?
Generally yes for future crawls. You cannot retroactively remove content a bot already fetched before the block, if that content was incorporated into a training run. The decision is forward-looking, not retroactive.
Does a high crawl count mean I am about to be cited?
No. Crawling is a prerequisite for citation, not a predictor of it. Training crawlers in particular can request thousands of pages from a site that is never cited anywhere. Keep crawl counts and citation counts in separate columns.
What about bots from operators not in the table - xAI, ByteDance, regional crawlers?
The table is scoped to operators documented consistently enough to give a reliable per-bot recommendation. Emerging and under-documented crawlers are tracked but held to a lower confidence level until their behaviour and purpose are published.
Do CDNs handle any of this automatically?
Some do. Cloudflare offers one-click AI-bot blocking categories and verified-bot detection, which is worth using for the verification step. The per-bot allow/block decision still has to be made deliberately, because a default "block all AI" setting removes you from citations too.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.