Every quantified figure on this page is vendor-reported, not first-party: it comes from Cloudflare's network-level bot classification, reproduced through secondary aggregators — Partial on this site's provenance scale. On that data, AI-related bots account for roughly a quarter of verified bot traffic, and some AI crawlers take over a thousand pages for every visitor they send back, versus under five for Google, a gap Cloudflare documented in its own crawl-to-click analysis.
Which figures here are first-party, and which are not
This site runs its own server-log studies, and it also reproduces vendor figures. Those are two very different kinds of evidence, and mixing them is the most common failure in crawler coverage. So they are kept in separate sections below, and neither borrows the other's credibility.
| Section | Evidence type | Population / denominator | Grade |
|---|---|---|---|
| What this site's own server logs show | First-party observation, one small site | Requests to this domain over a stated window | Traceable |
| Every ratio, share and composition figure below it | Vendor-reported, via secondary aggregators | Verified bot requests across Cloudflare's network — not the web | Partial |
| Blocking rates | Contested across three studies; no figure published here | Different domain samples in each study | Broken chain |
- This site has published no first-party crawl-to-referral ratio. Its own log studies measured which bots fetched a specific file, not how many pages a bot takes per visit it sends. Do not read the ratios below as anything this site measured.
- Every ratio here is a ratio of two counts over the same window: pages crawled by a bot, divided by referral visits attributed to that bot’s product. It is not a percentage.
- Verified bot traffic means requests a network operator confirmed came from a declared crawler by reverse-DNS or similar, the method described in Cloudflare’s 2025 breakdown. Spoofed and unverified requests are excluded.
- The denominator for every vendor figure is Cloudflare’s network, a large but non-random slice of the web. Sites not behind that network are invisible to it.
- None of these figures were independently re-derived for this page. The provenance audit explains why stitched secondary series are the norm in this field rather than the exception.
What this site's own server logs show
Fact The one crawler observation this site has made directly, and can stand behind, comes
from the 90-day llms.txt log test: after deploying an
llms.txt file and leaving it untouched, the file was requested a handful of times over three
months, overwhelmingly by SEO audit tools and generic scanners, and essentially never by the AI retrieval
bots that decide citations. That is a count of requests to one file on one site — not a ratio, not a
market share, and not generalisable.
The zero-to-cited log records this site's own crawl and citation events as they happen, and the crawl-to-citation latency study is the protocol designed to measure the gap between a first crawl and a first citation. Neither yet publishes a crawl-to-referral ratio. When one exists, it will appear in this section, not the next one.
What vendors report: the headline numbers
Everything from here to the methodology section is vendor-reported. Read each figure with its population and date attached.
of verified bot requests across Cloudflare's network were AI-related — training crawlers plus AI-search retrieval bots combined.
pages ClaudeBot crawled network-wide for every referral visit attributed to Claude.
of AI-crawler requests on that network taken by ClaudeBot alone — nearly double GPTBot's share.
How many pages does each crawler take per visit it sends?
Read plainly: Mistral's crawler took roughly 3,389 pages for every visitor it referred back, the most extractive ratio recorded among major operators in that data. ClaudeBot and GPTBot follow, at 2,237:1 and 217:1. Google's traditional crawler, included for contrast, sits at 4.6:1 — nearly three orders of magnitude apart. Trade coverage of the same dataset notes that Googlebot still out-crawls every AI operator in absolute volume, which is a separate point from the ratio. Anthropic documents ClaudeBot's purpose and opt-out separately from its answer-time web search tool.
Has the ratio moved over the past year?
That jagged note is deliberate. Secondary coverage of Cloudflare data reports figures at different points using seemingly different methodologies, and reconciling them into one clean line would misrepresent how confident the underlying picture is. A first-party, consistently-measured version is a planned deliverable of the Citation Index's crawler-monitoring component.
Which crawler takes the largest share of requests?
- ClaudeBot 16.28
- GPTBot 9.74
- Other AI crawlers 74
The year-over-year swing matters more than the snapshot. Cloudflare's own year in review and independent crawler inventories both show relative share moving several points within twelve months, so any ranking quoted today should be assumed temporary rather than structural.
How much of all bot traffic is AI?
- Traditional crawlers (Googlebot, Bingbot) 73.3
- AI training crawlers 20.3
- AI search / retrieval bots 6.5
How many sites block AI crawlers? (no figure published here)
A rising share of top web domains block at least one AI crawler in robots.txt, but independent counts from HTTP Archive-based analysis, a dedicated blocking report and a publisher-focused study all land in different places, because each counts different crawlers across a different domain sample. Broken chain We are not publishing a blocking-rate figure until one can be verified against a stated sample; the study designed to close that gap is the robots.txt AI-blocking census.
Blocking the wrong bot is the most common error here. GPTBot vs OAI-SearchBot covers the training-versus-retrieval distinction, and Google-Extended raises the same question on Google's side.
Why is the ratio so lopsided?
Three structural reasons, none of which require assuming bad faith:
- Different business models. A search engine's product is the click — crawl efficiency against referral volume has been optimized for two decades. An AI model's product is the answer itself.
- Training crawls are bulk operations. A crawler building a training corpus has no per-query reason to convert a crawl into a visit; the notion of a referral does not apply to that use case.
- Retrieval bots cite rather than redirect. Even bots fetching live to answer a question typically show a citation link, which is associated with a far lower click-through rate than a blue link — see AI search CTR statistics and the zero-click picture.
- 1 Training crawl Bulk corpus building. GPTBot, ClaudeBot, Google-Extended. No per-query notion of a referral exists.
- 2 Index or cache Content is stored for later retrieval. Nothing is visible to you at this stage.
- 3 Answer-time fetch A retrieval bot — OAI-SearchBot, PerplexityBot, Claude-SearchBot — pulls the page to answer a live question.
- 4 Citation The page may be quoted with a link. Most crawls never reach this step, which is what the ratio measures.
- 5 Referral visit A reader clicks the citation. The rarest outcome, and the denominator of every figure on this page.
What the crawl-to-referral ratio does not measure
Not citation. A page can be cited with no click following. The referral half counts visits, not mentions; the concordance study measures the mention side directly.
Not value. A visit from an AI answer may convert better or worse than a search visit — see the conversion benchmarks. The ratio treats every referral as equivalent, which no business does.
Not intent. A fetch for training and a fetch to answer a live question both count as one request, with different consequences for you.
Not cost. A thousand fetches of a cached static page is not comparable to a thousand database-backed renders, and a JavaScript-rendered page costs a crawler more — roughly nine times more crawl time in Google's case. Whether AI crawlers render at all is a separate open question.
Not reliable attribution. Referral detection depends on a referrer being visible at all. Some AI-surface traffic arrives without one and lands in direct traffic instead.
Hypothesis That last point makes us suspect published ratios are systematically overstated at the referral end. We cannot quantify by how much, and we will not guess a correction factor.
How to verify a bot is who it claims to be
Fact A user-agent string is a header the client chooses. Nothing authenticates it. Any script can present itself as any crawler, and Cloudflare has published evidence of undeclared crawlers evading no-crawl directives — exactly the failure mode a user-agent-only count cannot see. Take the requesting address from the log line, check it against the operator's published list or reverse-lookup rule, and record verified and unverified counts separately. Do not discard the unverified traffic: its size tells you how much of your apparent AI crawl load is something else.
Open question No public, disclosed-method figure exists for what share of self-declared AI crawler traffic fails verification across the web. We looked and did not find one, and that absence is notable given how often raw user-agent counts are published as fact. It is registered in the null results registry.
Seven ways your own log analysis will go wrong
Each of these changes a number without producing an error.
| Pitfall | What it does to the number |
|---|---|
| Matching on substrings | A loose pattern catches another product from the same operator. Anchor to the operator's published strings — OpenAI's and Perplexity's are authoritative. |
| Counting non-page requests | Images, stylesheets and scripts inflate request counts dramatically. Filter to HTML responses. |
| Ignoring status codes | A crawler hammering a redirect chain or a wall of 404s generates requests that mean something different from successful fetches. |
| Timezone drift | Logs and analytics often use different clocks. A misaligned window shifts both halves of the ratio. |
| CDN caching | If your edge serves a bot without touching the origin, origin logs undercount. Measure at one layer consistently. |
| Sampled analytics | A sampled referral count against an unsampled request count is not a ratio of anything. |
| Short windows | Crawl activity is bursty. A week can look nothing like the month around it, in either direction. |
A worked example against your own logs (hypothetical numbers)
The numbers in this example are invented to show the arithmetic. They are not observations from this site or any other.
- Pull a month of server logs and filter to HTML responses from a verified AI bot — say ClaudeBot makes 45,000 such requests.
- Cross-reference analytics for referral visits attributable to that product in the same window — say 12. Referrer strings differ by product, so check ChatGPT, Perplexity and Gemini separately rather than as one bucket.
- Compute the ratio: 45,000 ÷ 12 ≈ 3,750:1 — higher than the 2,237:1 network figure above, which would suggest this hypothetical site converts crawl volume into referrals less efficiently than the network average.
- Treat the gap as a prompt, not an instruction. A ratio worse than the published baseline is a reason to investigate content freshness, structure or category visibility — not an automatic reason to block.
The measurement standard writes this procedure down so anyone can copy it, and the crawl-side half of the Citation Index's panel program is designed to build a properly segmented baseline rather than one industry-wide average.
Does the ratio vary by industry or site type?
Open question A plausible mechanism exists — a reference-heavy site offers a crawler more distinct pages per session than a small, rarely-updated one — but the closest public attempt is a 30-day single-site log study, which is one site, not a vertical breakdown. No published, disclosed-method breakdown by vertical was located during this page's research.
What does not transfer between these numbers
| Does not transfer | Why not |
|---|---|
| A network-wide ratio to your site | Your content type, size and audience differ from the aggregate. The average describes a population you may not belong to. |
| One crawler's behaviour to another | Volumes and purposes differ by operator. The spread across crawlers on this page is itself the evidence. |
| This quarter's figure to next quarter's | These numbers have already reordered within a year. Recency matters more here than in most metrics. |
| Crawl volume to citation frequency | Being fetched often is not being quoted often. No published mapping between the two exists that we could verify. |
| Cloudflare-network figures to the whole web | The sample is large and not random. Sites not behind that network are invisible to it, and sites behind authentication are never reached at all. |
| Crawler behaviour to agentic browser behaviour | A browser agent acting for one user is a different traffic class from a bulk crawler — see what agentic browsers fetch. |
| Bot statistics to market share | Crawl volume reflects an operator's indexing appetite, not how many people use its product. Market share is a different measurement entirely. |
Common misreadings of these statistics
"AI crawlers are taking over bot traffic." The composition figure does not support that: traditional crawlers still account for the larger share in the data shown.
"A big ratio means block it." The ratio describes traffic economics. Blocking decisions also involve citation value, licensing position and reader reach. One input, not a conclusion.
"These are exact figures." They are network-derived estimates reproduced through secondary coverage. Carry the grade with the number.
"The trend line predicts next year." A series that has already reversed once is not a basis for extrapolation. We publish the points and decline the forecast.
"My site should match the average." Averages across a heterogeneous population describe almost nobody exactly.
What blocking, allowing and rate-limiting each cost
Blocking gives up citation and any referrals it brings, including the unlinked brand mentions that appear to correlate more strongly with citation than links do. You may still be described from other sources with less control, often via Wikipedia or via Reddit — and you lose the visibility your logs currently give you.
Allowing costs bandwidth and origin load, often smaller than assumed for cached static content, plus a strategic cost if licensing matters to your business.
Rate-limiting reduces load without a full block, but a crawler that repeatedly hits limits may simply fetch you less over time — a soft block you did not decide on.
The most expensive mistake is acting on unverified user-agent counts. Fix verification before any policy change. Where crawler access sits in the wider checklist is covered in the technical GEO audit.
Open questions these figures cannot answer
Open question How much crawl volume is duplicated across operators fetching the same pages? Nothing public separates unique content coverage from repeat fetching.
Open question What share of AI-driven visits arrive without a detectable referrer? This is the largest single uncertainty in every ratio on the page.
Open question Does crawl frequency respond to content updates, or run on a fixed schedule? A first-party log study could answer this and we have not seen one published.
Open question Are these ratios stable within a site over time, or do they swing as much as the network aggregate has? Site-level series long enough to tell are rare in public.
Null results we would publish
- A direct pull from the primary source contradicting these figures. We would publish the discrepancy and correct the page, even though it means retracting numbers already circulated.
- Verification showing most AI crawler traffic is unverified. Much of the published discourse, including parts of this page, would then rest on a weaker base than assumed.
- Site-level ratios clustering tightly around the network average. Our argument about non-generalisability would weaken, and we would say so.
- No relationship between crawl volume and citation — a question the citation half-life study and the tactic scoreboard both bear on. The practical case for tracking these numbers at all would shrink considerably.
What would change this page
- A direct pull from the primary data source, replacing secondary coverage and upgrading the provenance grade.
- Any published breakdown by vertical with a disclosed method, which would replace this page's stated gap with a measurement.
- A disclosed-method estimate of user-agent spoofing rates, which would change how every raw count here should be read.
- An operator publishing its own crawl and referral figures, allowing a check against network-derived estimates.
- Evidence that referral attribution misses a large share of AI-driven visits, which would revise every ratio downward.
Methodology & verification status
To see whether the crawling is turning into visibility: run your URLs through the AI Overview Exposure Checker. To follow the first-party work that would replace the vendor figures above: the studies index lists every log study on the roadmap, including the blocking census this page declines to pre-empt.
Namdev, R. (2026). AI crawler statistics (v3). Retrieved from https://ritiknamdev.com/blog/ai-crawler-statistics Published under CC BY 4.0 — reuse freely with attribution.
For what each crawler is actually for, see the AI Bot Registry. For the single most-confused pair in this data, see GPTBot vs OAI-SearchBot.