Statistics · Updated periodically

AI crawler statistics

How much of the web's bot traffic is AI, which crawlers take the most relative to what they give back, and how that's shifted year over year — with a verification status on every figure.

Every number here carries a verification status. Where the underlying methodology isn't public, that's stated next to the figure rather than presented as settled fact.

Ritik Namdev Ritik Namdev ·Published September 2026 ·Curated + verified ·14 min read ·Last verified September 2026
The short version

Every quantified figure on this page is vendor-reported, not first-party: it comes from Cloudflare's network-level bot classification, reproduced through secondary aggregators — Partial on this site's provenance scale. On that data, AI-related bots account for roughly a quarter of verified bot traffic, and some AI crawlers take over a thousand pages for every visitor they send back, versus under five for Google, a gap Cloudflare documented in its own crawl-to-click analysis.

Which figures here are first-party, and which are not

This site runs its own server-log studies, and it also reproduces vendor figures. Those are two very different kinds of evidence, and mixing them is the most common failure in crawler coverage. So they are kept in separate sections below, and neither borrows the other's credibility.

SectionEvidence typePopulation / denominatorGrade
What this site's own server logs showFirst-party observation, one small siteRequests to this domain over a stated windowTraceable
Every ratio, share and composition figure below itVendor-reported, via secondary aggregatorsVerified bot requests across Cloudflare's network — not the webPartial
Blocking ratesContested across three studies; no figure published hereDifferent domain samples in each studyBroken chain
What this page establishes
  • This site has published no first-party crawl-to-referral ratio. Its own log studies measured which bots fetched a specific file, not how many pages a bot takes per visit it sends. Do not read the ratios below as anything this site measured.
  • Every ratio here is a ratio of two counts over the same window: pages crawled by a bot, divided by referral visits attributed to that bot’s product. It is not a percentage.
  • Verified bot traffic means requests a network operator confirmed came from a declared crawler by reverse-DNS or similar, the method described in Cloudflare’s 2025 breakdown. Spoofed and unverified requests are excluded.
  • The denominator for every vendor figure is Cloudflare’s network, a large but non-random slice of the web. Sites not behind that network are invisible to it.
  • None of these figures were independently re-derived for this page. The provenance audit explains why stitched secondary series are the norm in this field rather than the exception.

What this site's own server logs show

Fact The one crawler observation this site has made directly, and can stand behind, comes from the 90-day llms.txt log test: after deploying an llms.txt file and leaving it untouched, the file was requested a handful of times over three months, overwhelmingly by SEO audit tools and generic scanners, and essentially never by the AI retrieval bots that decide citations. That is a count of requests to one file on one site — not a ratio, not a market share, and not generalisable.

The zero-to-cited log records this site's own crawl and citation events as they happen, and the crawl-to-citation latency study is the protocol designed to measure the gap between a first crawl and a first citation. Neither yet publishes a crawl-to-referral ratio. When one exists, it will appear in this section, not the next one.

What vendors report: the headline numbers

Everything from here to the methodology section is vendor-reported. Read each figure with its population and date attached.

26.7%

of verified bot requests across Cloudflare's network were AI-related — training crawlers plus AI-search retrieval bots combined.

Cloudflare-derived, via aggregators · May 2026
2,237:1

pages ClaudeBot crawled network-wide for every referral visit attributed to Claude.

Cloudflare-derived, via aggregators · Jul 2026
16.3%

of AI-crawler requests on that network taken by ClaudeBot alone — nearly double GPTBot's share.

Cloudflare-derived, via aggregators · Jul 2026

How many pages does each crawler take per visit it sends?

Read plainly: Mistral's crawler took roughly 3,389 pages for every visitor it referred back, the most extractive ratio recorded among major operators in that data. ClaudeBot and GPTBot follow, at 2,237:1 and 217:1. Google's traditional crawler, included for contrast, sits at 4.6:1 — nearly three orders of magnitude apart. Trade coverage of the same dataset notes that Googlebot still out-crawls every AI operator in absolute volume, which is a separate point from the ratio. Anthropic documents ClaudeBot's purpose and opt-out separately from its answer-time web search tool.

Has the ratio moved over the past year?

ClaudeBot's share of AI-crawler requests, Jul 2025 – Jul 2026 (illustrative trend line)
Jul '25Oct '25Jan '26Apr '26Jul '26
Connects reported endpoints (Jul 2025 ~13.2%, Jul 2026 16.3% per one source cited above, alongside a separately-reported single-month 9.74% figure) with an illustrative curve between them. The intermediate months are not independently confirmed data points and the discontinuity in the final segment reflects two different reporting methodologies rather than a real single-month drop — a limitation of stitching together secondary sources rather than a first-party time series.

That jagged note is deliberate. Secondary coverage of Cloudflare data reports figures at different points using seemingly different methodologies, and reconciling them into one clean line would misrepresent how confident the underlying picture is. A first-party, consistently-measured version is a planned deliverable of the Citation Index's crawler-monitoring component.

Which crawler takes the largest share of requests?

Share of AI-and-adjacent crawler requests, Cloudflare network, July 2026
  • ClaudeBot 16.28
  • GPTBot 9.74
  • Other AI crawlers 74
Source: Cloudflare-derived, via secondary aggregators. Population: AI-and-adjacent crawler requests on that network. A year earlier (Jul 2025), GPTBot led at ~13.2% and ClaudeBot trailed at ~11.2% — the two have since swapped relative positions.

The year-over-year swing matters more than the snapshot. Cloudflare's own year in review and independent crawler inventories both show relative share moving several points within twelve months, so any ranking quoted today should be assumed temporary rather than structural.

How much of all bot traffic is AI?

Composition of verified bot traffic, Cloudflare network, May 2026
  • Traditional crawlers (Googlebot, Bingbot) 73.3
  • AI training crawlers 20.3
  • AI search / retrieval bots 6.5
Source: Cloudflare-derived, via secondary aggregators. Population: verified bot requests across Cloudflare's network. 'AI training crawlers' and 'AI search / retrieval bots' are reported as separate categories by the underlying source; summed here for the 26.7% headline figure above.

How many sites block AI crawlers? (no figure published here)

A rising share of top web domains block at least one AI crawler in robots.txt, but independent counts from HTTP Archive-based analysis, a dedicated blocking report and a publisher-focused study all land in different places, because each counts different crawlers across a different domain sample. Broken chain We are not publishing a blocking-rate figure until one can be verified against a stated sample; the study designed to close that gap is the robots.txt AI-blocking census.

Blocking the wrong bot is the most common error here. GPTBot vs OAI-SearchBot covers the training-versus-retrieval distinction, and Google-Extended raises the same question on Google's side.

Why is the ratio so lopsided?

Three structural reasons, none of which require assuming bad faith:

  • Different business models. A search engine's product is the click — crawl efficiency against referral volume has been optimized for two decades. An AI model's product is the answer itself.
  • Training crawls are bulk operations. A crawler building a training corpus has no per-query reason to convert a crawl into a visit; the notion of a referral does not apply to that use case.
  • Retrieval bots cite rather than redirect. Even bots fetching live to answer a question typically show a citation link, which is associated with a far lower click-through rate than a blue link — see AI search CTR statistics and the zero-click picture.
Where a crawl request comes from, and what it can become
  1. 1 Training crawl Bulk corpus building. GPTBot, ClaudeBot, Google-Extended. No per-query notion of a referral exists.
  2. 2 Index or cache Content is stored for later retrieval. Nothing is visible to you at this stage.
  3. 3 Answer-time fetch A retrieval bot — OAI-SearchBot, PerplexityBot, Claude-SearchBot — pulls the page to answer a live question.
  4. 4 Citation The page may be quoted with a link. Most crawls never reach this step, which is what the ratio measures.
  5. 5 Referral visit A reader clicks the citation. The rarest outcome, and the denominator of every figure on this page.

What the crawl-to-referral ratio does not measure

Not citation. A page can be cited with no click following. The referral half counts visits, not mentions; the concordance study measures the mention side directly.

Not value. A visit from an AI answer may convert better or worse than a search visit — see the conversion benchmarks. The ratio treats every referral as equivalent, which no business does.

Not intent. A fetch for training and a fetch to answer a live question both count as one request, with different consequences for you.

Not cost. A thousand fetches of a cached static page is not comparable to a thousand database-backed renders, and a JavaScript-rendered page costs a crawler more — roughly nine times more crawl time in Google's case. Whether AI crawlers render at all is a separate open question.

Not reliable attribution. Referral detection depends on a referrer being visible at all. Some AI-surface traffic arrives without one and lands in direct traffic instead.

Hypothesis That last point makes us suspect published ratios are systematically overstated at the referral end. We cannot quantify by how much, and we will not guess a correction factor.

How to verify a bot is who it claims to be

Fact A user-agent string is a header the client chooses. Nothing authenticates it. Any script can present itself as any crawler, and Cloudflare has published evidence of undeclared crawlers evading no-crawl directives — exactly the failure mode a user-agent-only count cannot see. Take the requesting address from the log line, check it against the operator's published list or reverse-lookup rule, and record verified and unverified counts separately. Do not discard the unverified traffic: its size tells you how much of your apparent AI crawl load is something else.

Open question No public, disclosed-method figure exists for what share of self-declared AI crawler traffic fails verification across the web. We looked and did not find one, and that absence is notable given how often raw user-agent counts are published as fact. It is registered in the null results registry.

Match the full tokenAnchor user-agent patterns; substrings catch the wrong product from the same operator.
Verify the addressCheck the requesting IP against the operator published range or reverse-lookup rule.
Keep two seriesVerified and unverified counts, separately, from day one.
Never publish a raw UA countA user-agent header is self-reported text. Unverified counts are an upper bound, not a measurement.

Seven ways your own log analysis will go wrong

Each of these changes a number without producing an error.

PitfallWhat it does to the number
Matching on substringsA loose pattern catches another product from the same operator. Anchor to the operator's published strings — OpenAI's and Perplexity's are authoritative.
Counting non-page requestsImages, stylesheets and scripts inflate request counts dramatically. Filter to HTML responses.
Ignoring status codesA crawler hammering a redirect chain or a wall of 404s generates requests that mean something different from successful fetches.
Timezone driftLogs and analytics often use different clocks. A misaligned window shifts both halves of the ratio.
CDN cachingIf your edge serves a bot without touching the origin, origin logs undercount. Measure at one layer consistently.
Sampled analyticsA sampled referral count against an unsampled request count is not a ratio of anything.
Short windowsCrawl activity is bursty. A week can look nothing like the month around it, in either direction.

A worked example against your own logs (hypothetical numbers)

The numbers in this example are invented to show the arithmetic. They are not observations from this site or any other.

  1. Pull a month of server logs and filter to HTML responses from a verified AI bot — say ClaudeBot makes 45,000 such requests.
  2. Cross-reference analytics for referral visits attributable to that product in the same window — say 12. Referrer strings differ by product, so check ChatGPT, Perplexity and Gemini separately rather than as one bucket.
  3. Compute the ratio: 45,000 ÷ 12 ≈ 3,750:1 — higher than the 2,237:1 network figure above, which would suggest this hypothetical site converts crawl volume into referrals less efficiently than the network average.
  4. Treat the gap as a prompt, not an instruction. A ratio worse than the published baseline is a reason to investigate content freshness, structure or category visibility — not an automatic reason to block.

The measurement standard writes this procedure down so anyone can copy it, and the crawl-side half of the Citation Index's panel program is designed to build a properly segmented baseline rather than one industry-wide average.

Does the ratio vary by industry or site type?

Open question A plausible mechanism exists — a reference-heavy site offers a crawler more distinct pages per session than a small, rarely-updated one — but the closest public attempt is a 30-day single-site log study, which is one site, not a vertical breakdown. No published, disclosed-method breakdown by vertical was located during this page's research.

What does not transfer between these numbers

Does not transferWhy not
A network-wide ratio to your siteYour content type, size and audience differ from the aggregate. The average describes a population you may not belong to.
One crawler's behaviour to anotherVolumes and purposes differ by operator. The spread across crawlers on this page is itself the evidence.
This quarter's figure to next quarter'sThese numbers have already reordered within a year. Recency matters more here than in most metrics.
Crawl volume to citation frequencyBeing fetched often is not being quoted often. No published mapping between the two exists that we could verify.
Cloudflare-network figures to the whole webThe sample is large and not random. Sites not behind that network are invisible to it, and sites behind authentication are never reached at all.
Crawler behaviour to agentic browser behaviourA browser agent acting for one user is a different traffic class from a bulk crawler — see what agentic browsers fetch.
Bot statistics to market shareCrawl volume reflects an operator's indexing appetite, not how many people use its product. Market share is a different measurement entirely.

Common misreadings of these statistics

"AI crawlers are taking over bot traffic." The composition figure does not support that: traditional crawlers still account for the larger share in the data shown.

"A big ratio means block it." The ratio describes traffic economics. Blocking decisions also involve citation value, licensing position and reader reach. One input, not a conclusion.

"These are exact figures." They are network-derived estimates reproduced through secondary coverage. Carry the grade with the number.

"The trend line predicts next year." A series that has already reversed once is not a basis for extrapolation. We publish the points and decline the forecast.

"My site should match the average." Averages across a heterogeneous population describe almost nobody exactly.

What blocking, allowing and rate-limiting each cost

Blocking gives up citation and any referrals it brings, including the unlinked brand mentions that appear to correlate more strongly with citation than links do. You may still be described from other sources with less control, often via Wikipedia or via Reddit — and you lose the visibility your logs currently give you.

Allowing costs bandwidth and origin load, often smaller than assumed for cached static content, plus a strategic cost if licensing matters to your business.

Rate-limiting reduces load without a full block, but a crawler that repeatedly hits limits may simply fetch you less over time — a soft block you did not decide on.

The most expensive mistake is acting on unverified user-agent counts. Fix verification before any policy change. Where crawler access sits in the wider checklist is covered in the technical GEO audit.

Open questions these figures cannot answer

Open question How much crawl volume is duplicated across operators fetching the same pages? Nothing public separates unique content coverage from repeat fetching.

Open question What share of AI-driven visits arrive without a detectable referrer? This is the largest single uncertainty in every ratio on the page.

Open question Does crawl frequency respond to content updates, or run on a fixed schedule? A first-party log study could answer this and we have not seen one published.

Open question Are these ratios stable within a site over time, or do they swing as much as the network aggregate has? Site-level series long enough to tell are rare in public.

Null results we would publish

  • A direct pull from the primary source contradicting these figures. We would publish the discrepancy and correct the page, even though it means retracting numbers already circulated.
  • Verification showing most AI crawler traffic is unverified. Much of the published discourse, including parts of this page, would then rest on a weaker base than assumed.
  • Site-level ratios clustering tightly around the network average. Our argument about non-generalisability would weaken, and we would say so.
  • No relationship between crawl volume and citation — a question the citation half-life study and the tactic scoreboard both bear on. The practical case for tracking these numbers at all would shrink considerably.

What would change this page

  • A direct pull from the primary data source, replacing secondary coverage and upgrading the provenance grade.
  • Any published breakdown by vertical with a disclosed method, which would replace this page's stated gap with a measurement.
  • A disclosed-method estimate of user-agent spoofing rates, which would change how every raw count here should be read.
  • An operator publishing its own crawl and referral figures, allowing a check against network-derived estimates.
  • Evidence that referral attribution misses a large share of AI-driven visits, which would revise every ratio downward.

Methodology & verification status

First-party dataOne observation only, from the 90-day llms.txt log test on this site. No first-party crawl-to-referral ratio has been published here. Traceable
Primary source (vendor)Cloudflare Radar, which classifies bot traffic across its network at scale.
This page's sourcingVendor figures reproduced via secondary aggregator coverage of Cloudflare data, not pulled directly from Radar for this page. Partial
Known limitationCloudflare's network reflects sites behind Cloudflare, a large but non-random sample of the web. Figures may not generalize to unprotected origin infrastructure, and content behind authentication or reached via an MCP server is not described by any log-based ratio.
Next stepA future edition will pull directly from Radar's public API rather than secondary coverage, upgrading these figures toward Traceable .
Where to go next

To see whether the crawling is turning into visibility: run your URLs through the AI Overview Exposure Checker. To follow the first-party work that would replace the vendor figures above: the studies index lists every log study on the roadmap, including the blocking census this page declines to pre-empt.

How to cite this
Namdev, R. (2026). AI crawler statistics (v3). Retrieved from https://ritiknamdev.com/blog/ai-crawler-statistics

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

For what each crawler is actually for, see the AI Bot Registry. For the single most-confused pair in this data, see GPTBot vs OAI-SearchBot.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Cloudflare Radar — AI crawler and bot trafficradar.cloudflare.com InfoQ — Cloudflare 2025 AI bots report coverageinfoq.com/news/2025/12/cloudflare-2025-ai-bots OpenAI — GPTBot documentationplatform.openai.com/docs/bots OpenAI — Bots and crawler reference (developer docs)developers.openai.com/api/docs/bots Cloudflare — From Googlebot to GPTBot: who is crawling your site in 2025blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Cloudflare — Crawlers, clicks and AI bots: the training-to-referral gapblog.cloudflare.com/crawlers-click-ai-bots-training Cloudflare — Radar 2025 Year in Reviewblog.cloudflare.com/radar-2025-year-in-review Cloudflare — Perplexity is using stealth, undeclared crawlers to evade no-crawl directivesblog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives Search Engine Journal — Cloudflare report: Googlebot tops AI crawler trafficwww.searchenginejournal.com/cloudflare-report-googlebot-tops-ai-crawler-traffic/563303 Seomator — Crawl-to-refer ratio for AI crawlers and LLM botsseomator.com/blog/crawl-to-refer-ratio-ai-crawlers-llm-bots Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots Paul Calvano — AI bots and robots.txtpaulcalvano.com/2025-08-21-ai-bots-and-robots-txt Technology Checker — robots.txt AI crawler blocking reporttechnologychecker.io/blog/robots-txt-ai-crawlers-blocking-report BuzzStream — Publishers blocking AI crawlers studywww.buzzstream.com/blog/publishers-block-ai-study Perplexity — Crawler documentationdocs.perplexity.ai/docs/resources/perplexity-crawlers Anthropic Support — Does Anthropic crawl the web, and how site owners can block the crawlersupport.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Anthropic — Web search tool documentationplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Google Search Central — Overview of Google crawlers and fetchersdevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers Google Search Central — Managing crawl budget for large sitesdevelopers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget AmICited — Google-Extended: what it does and whether to block itwww.amicited.com/blog/google-extended-what-it-does-should-you-block-it Digital Applied — Agentic crawler behaviour: a 30-day site log studywww.digitalapplied.com/blog/agentic-crawler-behavior-30-day-site-log-study Onely — Google needs 9x more time to crawl JavaScript than HTMLwww.onely.com/blog/google-needs-9x-more-time-to-crawl-js-than-html Web Almanac 2025 — Performance chapteralmanac.httparchive.org/en/2025/performance ritiknamdev.com — Where AI SEO statistics come fromritiknamdev.com/blog/where-ai-seo-statistics-come-from
FAQ

Frequently asked questions

Where does this data actually come from?
The ratios are derived from Cloudflare's network-level bot classification, reported via secondary aggregators rather than pulled directly from Cloudflare Radar for this page. That distinction is stated plainly in the methodology section. Treat these as directionally reliable, not lab-grade precise.
Is a low crawl-to-referral ratio always better?
It means that engine sends you more visitors per page it takes, which is favorable if your goal is referral traffic. But a high ratio from an engine that cites you prominently in its answers may still be worth allowing. The ratio measures traffic economics, not citation value.
Why is Google's ratio so much lower than the AI-specific crawlers?
Googlebot has spent 25+ years optimizing crawl efficiency against a search product with billions of daily referrals. The AI crawlers are optimizing for model training and answer quality, where a referral back to the source isn't the product's primary goal, the way it is for a search engine's click-through business model.
Will these numbers stay accurate?
No. This is one of the fastest-moving figures in AI search, with year-over-year swings already observed. ClaudeBot and GPTBot swapped relative rank between July 2025 and July 2026. Treat any specific ratio as a snapshot, dated as shown.
Can I trust the user-agent string in my own logs?
Only partly. A user-agent header is self-reported text and anyone can send any value. Some traffic claiming to be a major AI crawler is not. Verifying the requesting address against the operator's published ranges is the standard defence, and it changes the numbers on most sites that try it.
Does a high crawl-to-referral ratio mean an engine is behaving badly?
Not by itself. A high ratio can reflect a product that answers questions without needing to send a click, which is a design choice rather than misconduct. It can also reflect a crawler that is inefficient. The ratio alone cannot distinguish those, and treating it as a verdict overstates what it measures.
Why does this page not publish a single combined AI crawler number?
Because a combined number would mix training crawlers with answer-time fetchers, and those have different purposes, different volumes and different consequences for a publisher. Any aggregate hides the one distinction most likely to change what a site owner should do.
How often should I recompute my own figures?
Monthly is enough for most sites, and quarterly is defensible for small ones. The underlying figures move fast, but so does the noise. Recomputing weekly usually produces variation you will over-interpret rather than signal you can act on.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.