Statistics · Claude — the least-measured surface

Claude citation statistics

Claude is the least-measured major AI search surface: no large-scale citation study has been published for it. This page is an honest accounting of that gap, and of the one configuration decision that is fully documented.

Ritik Namdev Ritik Namdev ·Published September 2026 ·Genuine research gap ·10 min read ·Last verified September 2026
The short version

Claude citation data barely exists. It is the least-measured of the major AI search surfaces: no vendor or researcher has published a large-scale study of what Claude cites, at any sample size, with any disclosed method. Everything below the bot map on this page is qualitative. The thing that would fill the gap is the AI Citation Index, where the Claude Citation Report is registered as the flagship first study — for exactly this reason.

What this page establishes
  • No quantified source-type breakdown for Claude exists in public. Every other engine page on this site cites a vendor study of millions of citations; for Claude that study has not been published. The absence is the finding.
  • Claude has a documented web search tool that performs live retrieval and attaches citations to search-derived claims. Graded fact, from Anthropic's own platform documentation.
  • Claude-SearchBot is the switch. It performs the live retrieval that produces a citation. ClaudeBot is the training crawler and has no effect on today's citations. This is the one part of the page where guidance is concrete.
  • Claude is reported to be more citation-conservative than ChatGPT — fewer, more verifiable sources rather than a high volume of loosely related ones. Graded hypothesis: consistent across the little coverage that exists, quantified by none of it.
  • Nothing here is a Claude citation rate, a source share, or a ranking-overlap figure, because none has been published. Borrowing one from ChatGPT or Perplexity would misrepresent a differently built system serving a differently composed audience.

Why this page is short on findings

Every other statistics page here — ChatGPT, Perplexity, AI Overviews — can cite a study analysing millions of citations, however partial its provenance. For Claude, that study does not exist in public. The closest published comparisons, from Discovered Labs, Leapd and Profound, are qualitative on Claude specifically.

We looked, and the absence itself is the finding — the same pattern the provenance audit documents across the field, and the reason the null results registry exists. This page will not paper over it with a number borrowed from a different platform, because that is the specific failure the site was built to interrupt. What it will do is be precise about what is known.

Which engines have public citation data, and which do not

EngineLarge-scale public citation study?Quantified source-type breakdown?Grade of the category on this site
ChatGPTYes — vendor studies with disclosed sample sizesYesPartial — see ChatGPT citation statistics
PerplexityYes — visible source lists make collection easiestYesPartial — see Perplexity citation statistics
Google AI OverviewsYes — including a repeated measure by one organisationDirectional only, vendor corpusTraceable on the overlap trend — see AI Overview statistics
Gemini (app)No — circulating figures describe a different Google surfaceNoBroken chain — see Gemini citation statistics
ClaudeNo — none published at any sample sizeNoQualitative only; this page

Read the second column as the reason this page is short. Claude and Gemini are the two surfaces with no published citation measurement, and they are unmeasured for different reasons: Gemini because coverage conflates it with two other Google products, Claude because nobody has run the study.

What is publicly known about Claude citations

Fact

Claude has a dedicated web search tool. It gives Claude live access to current web content beyond its training cutoff and attaches citations to search-derived claims. The wider Anthropic documentation is the only first-party account of how the tool behaves.

Fact

Newer web-search versions can filter results with code before they reach context. This dynamic-filtering capability is documented for Claude 4.6 and later, and it may affect which sources end up surfacing in a cited answer.

Hypothesis

Claude is reported to be more citation-conservative than ChatGPT — citing fewer, more verifiable sources rather than a high volume of loosely related ones, and leaning away from vague or promotional content where a better-cited alternative exists. That pattern is consistent across the few pieces of coverage we found. No quantified study backs it up, and nothing equivalent to ChatGPT's Wikipedia dependency or Perplexity's Reddit dependency has been measured here.

Claude is the least-measured major AI search surface. No large-scale citation study exists publicly for it — which is exactly why it's the first flagship study registered in the Citation Index.

Share on X

The Claude bot map, and the one switch that matters

Three bots, three jobs, and only one determines whether you can be cited. This is the section with genuinely actionable guidance, because the bot behaviour is documented even where the citation behaviour is not.

BotRoleRelevant to citation?Cost of blocking it
ClaudeBotTraining crawlNo — affects future models onlyZero, in citation terms
Claude-SearchBotLive retrieval for the web search toolYes — this is the one that mattersRemoval from Claude's answers, silently
Claude-UserFetches a URL a user explicitly referencesOnly for that specific referenced pageFrustrating your own readers
Allow Claude-SearchBotThe live retrieval bot. Blocking it removes you from Claude's answers entirely, with no error and no report.
Allow Claude-UserFetches only a page a person has explicitly pasted. Blocking it mostly frustrates your own readers.
Decide on ClaudeBot separatelyThe training crawler. A content-policy question. Blocking it costs nothing in citation terms.
Never write a blanket AI blockA single rule that catches all three is the most expensive mistake available here.

The mistake that costs most is a robots.txt rule intended to opt out of AI training that catches Claude-SearchBot alongside ClaudeBot. The owner believes they declined training. They also removed themselves from Claude's search answers, and nothing about that outcome is visible to them — no error message, no report, just an absence.

Verify by checking your server logs for each user agent individually, confirming by IP rather than trusting the user-agent string, which is trivially spoofed. Anthropic's own crawler and opt-out article is the first-party reference, and the bot registry holds the verification patterns. One footnote: older block lists reference anthropic-ai and Claude-Web, which are deprecated rather than active. They are harmless to leave in place, and they do not cover the current bots.

The same training-versus-retrieval split exists at OpenAI and Perplexity; GPTBot vs OAI-SearchBot is the same mistake in OpenAI's vocabulary, and the blocking census counts how often it gets made.

Why this gap exists, structurally

"Nobody has studied it" invites the question of why not, and the answer explains how this field's evidence gets produced. Almost all citation research comes from visibility vendors — not academics, not publishers, not the AI companies — who build citation-tracking infrastructure because they sell access to it.

Vendor coverage therefore follows commercial demand rather than research value. Customers ask about ChatGPT because it has the largest consumer footprint, and about Google because it is Google. A platform with a smaller, more technical user base generates fewer tickets asking "am I visible there." Later arrival compounds it: retrofitting a new surface into an existing measurement pipeline is real engineering work with unclear payback.

Nothing about the platform makes it hard to measure. Claude cites sources in its answers. A researcher could run a query set against it and record what comes back, exactly as for any other engine. The procedure is written down in the measurement standard, and what to record is in the dataset strategy. A genuine gap with no technical barrier is exactly the profile of a research opportunity, which is why this platform sits first in the Index roadmap despite being the smallest surface tracked.

What transfers from other engines, and what does not

Being explicit about which borrowings are defensible is more useful than a blanket "we don't know."

Likely transfers: content-level tactics with a mechanism. The Princeton GEO study's findings on quotations, statistics and cited sources rest on a mechanism — specific, attributable content is easier and safer to quote — that is not engine-specific. Any retrieval system composing an answer from sources faces the same problem. This is the most defensible borrowing available.

Likely transfers: access requirements. A bot that cannot fetch a page cannot cite it. Less a finding than a physical constraint, which is why the technical GEO audit and the JavaScript rendering question come before any content work. Whether an llms.txt file helps is settled: it does not appear to.

Probably does not transfer: source-type skew. ChatGPT's encyclopedic lean and Perplexity's community lean are very different from each other, and that variance is itself evidence the skew is engine-specific. Assuming Claude resembles either would be a guess dressed as an inference.

Probably does not transfer: ranking-overlap figures. Overlap with Google's top 10 runs from roughly 12% to roughly 38% across the engines that have been measured. Separate work found low overlap between the engines themselves, and Seer found 87% agreement between SearchGPT and Bing's top results. With that much spread, no single figure predicts an unmeasured engine. Citation density is unknown for the same reason.

Who uses Claude, and why it changes the question

Hypothesis Claude has meaningful adoption among developers and technical practitioners, partly through coding assistants and IDE integrations rather than the consumer chat interface — a very different composition from the consumer-weighted platforms behind the market share figures. It also sees professional and enterprise use in contexts where careful, well-sourced answers matter more than speed.

That has a specific implication. A user base weighted toward technical and professional questions asks different questions, so documentation, technical references and specialist writing plausibly matter more here than on a platform handling a broad consumer mix. Borrowing figures from a consumer-weighted platform is therefore doubly unsafe: the retrieval system is different and the query distribution feeding it likely differs too.

The same caution applies to the traffic-side figures. Referral volume, CTR and conversion benchmarks are all borrowed from a different audience here. This is reasoning from observable adoption patterns rather than citation data — a reason to investigate, not a conclusion to act on.

How to measure Claude yourself

The absence of published research has one upside: anyone can produce the first measurement of their own presence here, with no tooling beyond a spreadsheet. Build a fixed panel of fifteen to thirty questions your actual audience would ask, weighted toward the specific and technical. That is where the user base skews, and where a smaller site is most likely to appear at all.

Run each question more than once, because generated answers vary between runs. Record whether the search tool was used, which sources were cited, whether you were linked and whether you were named — those last two move independently and are worth different amounts. Hold conditions constant and write them down. Keep the raw records and repeat quarterly.

On the content side, the safest working guess is to apply the tactics with the strongest general evidence: clear citations, direct quotations, disclosed statistics. Put extra weight on making claims easy to verify, since that property keeps recurring in descriptions of Claude's source selection. It is a working guess, not a proven playbook; the version written for a platform that has been measured is how to get cited by ChatGPT.

What would change this page

A published source-mix study with a disclosed method, from anyone. A ranking-overlap measurement for Claude specifically. First-party citation reporting from Anthropic, which no AI company currently provides. Or our own Claude Citation Report, which is the reason this page exists in its current form. When any of those land, the page gets rebuilt rather than amended, and the version history stays visible. A page whose central claim is "nobody has measured this" should not quietly become a page with numbers on it, as though the gap had never been there.

Next step

Read the Citation Index roadmap — the Claude Citation Report is its flagship first study, and it is the thing that would replace every hypothesis on this page with a measurement. Results ship first through the newsletter.

Verification status

Everything on this page other than the bot map and the documented tool behaviour is graded Partial at best, and most of it is Hypothesis . No figure on this page is a Claude citation statistic, because none has been published. It will be substantially rewritten — not incrementally updated — once the first Claude Citation Report ships.

The crawl-side figures it will be read against come from Cloudflare's crawl-to-click analysis, discussed in the crawler statistics page. The study is listed on the studies index, the method is published in full, and the first first-party log studies published here are zero to cited and the llms.txt study.

How to cite this
Namdev, R. (2026). Claude citation statistics (v1). Retrieved from https://ritiknamdev.com/blog/claude-citation-statistics

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

This gap is the reason the Claude Citation Report is registered as the first flagship study in the AI Citation Index roadmap. See the AI Bot Registry for the full Claude bot map.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Claude Platform Docs — web search toolplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Anthropic — Documentation homedocs.anthropic.com Anthropic Support — Does Anthropic crawl the web, and how site owners can block the crawlersupport.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Discovered Labs — AI citation patterns: ChatGPT, Claude, Perplexitydiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Discovered Labs — How each platform cites sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Leapd — How ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Ziptie — How does ChatGPT choose its sources?ziptie.dev/blog/how-does-chatgpt-choose-its-sources Ziptie — How original research wins AI citationsziptie.dev/blog/how-original-research-wins-ai-citations Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors GEO: Generative Engine Optimization — Aggarwal et al., KDD 2024arxiv.org/abs/2311.09735 arXiv — GEO paper, full PDFarxiv.org/pdf/2311.09735 Ahrefs — How much do AI search engines overlap with each other?ahrefs.com/blog/ai-search-overlap Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Cloudflare — Crawlers, clicks and AI bots: the training-to-referral gapblog.cloudflare.com/crawlers-click-ai-bots-training InfoQ — Cloudflare 2025 AI bots report coverageinfoq.com/news/2025/12/cloudflare-2025-ai-bots Seomator — Crawl-to-refer ratio for AI crawlers and LLM botsseomator.com/blog/crawl-to-refer-ratio-ai-crawlers-llm-bots Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots OpenAI — Bots and crawler documentationplatform.openai.com/docs/bots Perplexity — Crawler documentationdocs.perplexity.ai/docs/resources/perplexity-crawlers Onely — Writing LLM-friendly contentwww.onely.com/blog/llm-friendly-content ritiknamdev.com — Where AI SEO statistics come fromritiknamdev.com/blog/where-ai-seo-statistics-come-from
FAQ

Frequently asked questions

Why is there so little public data on Claude's citations?
Claude's web search tool is newer than ChatGPT Search or Perplexity, and no major vendor has published a large-scale citation study for it. This is a genuine, verified research gap rather than us withholding data.
Does Claude prefer certain source types the way ChatGPT prefers Wikipedia?
It is reported to be more conservative and selective than ChatGPT, favouring well-verified, clearly-sourced content over vague or promotional pages. But no quantified source-type breakdown, like the ChatGPT or Perplexity figures, exists publicly.
What is the difference between ClaudeBot, Claude-SearchBot and Claude-User?
ClaudeBot trains future models. Claude-SearchBot performs live retrieval for Claude's web search tool and is the one that actually produces citations. Claude-User only fetches a page when a person explicitly references its URL in a conversation.
Does blocking ClaudeBot remove me from Claude's answers?
No, and this is the most consequential misunderstanding on this page. ClaudeBot is the training crawler. Claude-SearchBot performs the live retrieval that produces citations. Blocking the first is a content-policy decision with no effect on citation. Blocking the second removes you from Claude's answers.
Should I optimise for Claude differently than for other engines?
On current evidence, no, and anyone telling you otherwise is guessing with more confidence than the data supports. The defensible approach is applying the tactics with the strongest general evidence, with extra weight on making claims easy to verify, since that is the one property repeatedly associated with Claude's source selection.
Is Claude worth attention if nobody has measured it?
That depends on your audience rather than on the measurement. Claude has substantial usage among technical and professional users. If that describes your market, the absence of published research is a reason to measure it yourself, not a reason to ignore it.
Is this the platform you are most likely to research first with the Citation Index?
Yes. Claude is registered as the flagship first study in the Index roadmap. It is the least-measured major surface, so the research value per query is highest here.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.