Original research · Provenance audit

Where AI SEO statistics actually come from

We took twelve of the most-quoted numbers in AI search and followed each one back, hop by hop, to whoever measured it first.

Five hold up completely. Four name a source but withhold the method. Three dead-end with no primary data anywhere in the chain. And the single most-repeated figure in the field is three different studies wearing the same number.

Ritik Namdev Ritik Namdev ·Published September 2026 ·12 claims traced ·22 min read ·Last verified September 2026
The short answer

About 40% of the field's most-quoted statistics trace cleanly to a named study with a disclosed method. The rest either name a source that withholds its methodology, or dead-end entirely. The failure mode is not fabrication — it is that checkable and uncheckable numbers circulate in the same sentence with nothing to distinguish them.

The five provenance checks

These are the checks used to grade every figure in this audit, and they are the grading scheme applied to every statistic published anywhere on this site. Each takes about two minutes. Together they catch nearly everything in the ledger below.

  1. Click the citation, not the summary. Follow the actual link, one hop at a time, until you reach a primary source, a circular loop, or a page that cites nothing. If there is no link at all, stop there — you already have your answer.
  2. Find the population. A percentage without a denominator is not a statistic. "US searches" and "queries that trigger an AI Overview" are different worlds, and conflating them is how the 58% collision below happens.
  3. Find the date and the collection window. These products change monthly. An AI-search figure older than roughly six months may describe a product that no longer exists in that form, and an undated figure decays silently.
  4. Check the method is disclosed well enough to disagree with. Sample size, technique, and how the measurement was taken. The real test of disclosure is whether a critic could attack it — not whether an organisation's name appears.
  5. Ask what would falsify it. If no possible observation would contradict the claim, it is a position rather than a finding, however many decimal places it carries.

Grades follow directly from those checks. Traceable means the chain reaches a named study with sample, date and method disclosed. Partial means it reaches a named organisation whose underlying method or corpus is withheld. Broken chain means no primary source is reachable, or the sources contradict each other irreconcilably. And when a number passes all five, cite the primary source rather than the page you found it on — every extra hop is one more place the method can fall out.

Research status Results published

A completed provenance audit. Each figure below was traced by hand to its earliest locatable source; the method and the full result set are on this page.

What is known
  • Which of the field's most-quoted statistics trace cleanly to a named study with a disclosed method, and which do not.
  • Where each broken chain terminates — the specific point at which provenance stops.
  • The five checks used to grade a figure, published so anyone can apply them.
What is not yet known
  • Whether the untraceable figures are wrong — an untraceable number is unverifiable, not disproven.
  • Whether the traced sample is representative of every figure in circulation; it covers the most-quoted ones, not all of them.

An unchecked statistic is a liability in a specific way. When you quote a figure in a client deck or a strategy doc, you are implicitly vouching for what it measures. It may instead measure something adjacent — a different population, a different time window, a different definition of the same word. The recommendation built on top of it is then wrong in a way nobody can see until it fails.

  • A budget reallocated on "AI traffic converts 23× better" — one of three unreconciled vendor claims.
  • A crawler blocked on a crawl-to-referral ratio drawn from a different vertical entirely.
  • A journalist repeating "AI Overviews cut clicks by 58%", a figure this audit found conflates two separate studies.

None of these require bad faith. They require only that a plausible-sounding number circulated widely enough, for long enough, that checking it stopped feeling necessary.

So this audit asks one question of twelve claims: can a reader get from the statistic to the thing that produced it? Not "is it true" — that is a much harder question. Just: is the path there open.

An unchecked statistic is a liability. When you quote a figure, you vouch for what it measures — and if it measures something adjacent, the recommendation built on it is wrong in a way nobody can see until it fails.

Share on X

How we traced twelve claims

The method is deliberately dull, because the value is in doing it consistently rather than cleverly.

Tracing procedure, per claim
  1. 01 Collect the claim As phrased on a page that repeats it, verbatim
  2. 02 Follow the link Whatever the page cites, one hop at a time
  3. 03 Repeat to the end Until a primary source, a loop, or a dead end
  4. 04 Check the method Is sample, date and technique disclosed?
  5. 05 Grade and record Traceable / partial / broken, with the chain
Claim selection Twelve statistics chosen for how often they appear in AI-search writing, not for how likely they were to fail. Selection was judgement-based, which is the audit's main weakness — see limitations.
Chain following Each claim was followed one citation at a time, recording every intermediate page, until reaching a primary source, a circular loop, or a page that cited nothing.
Date All tracing performed in the first week of September 2026. Chains change as pages are edited, so this is a snapshot with a timestamp, not a permanent verdict.

What we found

Provenance grade — 12 claims traced, September 2026
  • Traceable to primary 5
  • Partial — source named, method withheld 4
  • Broken chain 3
First-party result of the audit described on this page. A small, judgement-selected sample: treat the proportions as indicative of a pattern, not as a population estimate.

The headline is less damning and more interesting than "SEO statistics are made up." Nearly half trace cleanly. The problem is the other half — and specifically that a reader has no way to tell which half a given number belongs to, because both are presented in the same confident register.

42%

of the twelve traced claims graded fully traceable to a named study with disclosed method and sample.

This audit · Sep 2026

Chain length is the practical signal. One or two hops is normal and healthy — a trade publication summarising a study is doing its job. Past three hops, the original method has almost always been stripped away, and what survives is a number and a vibe.

5 / 12

claims traced cleanly to a named study with a disclosed sample, method and date.

This audit · Sep 2026
3 / 12

claims where no primary source was reachable, or where sources contradict irreconcilably.

This audit · Sep 2026
3

distinct studies found circulating under the same '58%' figure, each measuring something different.

This audit · Sep 2026

The full ledger

Every claim, its grade, and where the chain actually ends. This is the part worth bookmarking.

Provenance ledger — traced September 2026
Claim as commonly quotedChain ends atMethod disclosed?Grade
"76% of AI Overview citations come from top-10 pages" — later "down to 38%" Ahrefs, Jul 2025 (1.9M citations) and Mar 2026 (863K SERPs) Yes — sample and window stated Traceable
"Only 12% of AI-cited URLs rank in Google's top 10" Ahrefs, 2026 Yes Traceable
"58.5% of US searches end without a click" SparkToro with Datos clickstream, 2024 Yes — panel and geography stated Traceable
"GEO tactics lift visibility up to 40%" Aggarwal et al., KDD 2024 — peer-reviewed, 10,000-query benchmark Yes — fully specified Traceable
"80% of consumers rely on zero-click results in 40%+ of searches" Bain & Company survey, Dec 2024 Yes — survey design stated Traceable
"AI crawlers take 2,237 pages per referral" (and similar ratios) Cloudflare network data — but usually quoted via aggregators, not Radar itself Partly — the derivation is rarely shown Partial
"ChatGPT's citations are 47.9% Wikipedia" Profound, from a reported 680M-citation corpus No — corpus and query set not public Partial
"Perplexity cites Reddit 46.7% of the time" Same corpus as above No Partial
"Only 11% of domains are cited by both ChatGPT and Perplexity" Vendor analysis; query set not published No Partial
"AI Overviews cut clicks by 58%" A different study from the 58.5% zero-click figure — routinely conflated with it Varies by retelling; the conflation destroys the meaning Broken chain
"AI traffic converts 4.4× better" / "23× better" / "14.2% vs 2.8%" Three separate vendor datasets that cannot be reconciled Partially, but definitions of "conversion" differ Broken chain
Round-number market-size and adoption forecasts Content-farm pages citing each other; no primary data located in the loop No Broken chain

A live thirteenth case, not yet in the ledger, shows the same method working on a fresh claim. "Google AI Mode fans a query out into 8–16 sub-queries" is repeated across dozens of SEO articles from 2025 and 2026, almost always without a citation. Google's documentation confirms the fan-out mechanism exists, and a granted patent (US11663201B2) describes eight distinct sub-query types the system may generate. Nowhere in that patent, or in any Google publication located during this research, is a specific per-query count of "8 to 16" stated. Graded Broken chain — plausible, oft-repeated, untraceable to any primary count. Full write-up at the Query Fan-Out Corpus study.

One trace, in full: how "2,237 pages per referral" was checked

Ledgers are abstract. One trace, step by step, exactly as it happened.

  1. The claim, as encountered: a marketing blog post stated "Anthropic's ClaudeBot crawls 2,237 pages for every visitor it sends back," with no link attached.
  2. Hop one: a search for the exact phrase surfaced a second aggregator site repeating the same number, this time with a link to a third-party SEO-tools blog.
  3. Hop two: that SEO-tools blog cited "Cloudflare data, July 2026" in prose, again without a direct link to Cloudflare Radar itself, and rounded the figure slightly differently (2,200-ish) than the first source.
  4. Hop three: Cloudflare Radar's own public dashboard was checked directly. It exposes bot-traffic classification and category-level breakdowns. The specific per-crawler ratio, as phrased, was not found on Radar as a standalone published figure. A third party most likely derived it from Radar's underlying data rather than quoting a Cloudflare publication.
  5. Verdict: graded Partial . The underlying network is almost certainly Cloudflare's, and the figure is plausible against the general pattern Cloudflare has described publicly. But nobody in the chain shows the derivation — which time window, which subset of traffic, how "referral" was defined. A reader repeating "2,237:1" is repeating a number nobody has shown their work for.

That is what "partial" looks like in practice: not a fabricated number, but one that has lost its method somewhere between the primary network data and the blog post that made it famous.

The 58% problem

This is the clearest finding in the audit, because it shows how a statistic can be simultaneously accurate and useless. Three separate, legitimate pieces of research produced numbers near 58%. They measure entirely different things. In circulation, they have collapsed into one figure people quote interchangeably. The CTR statistics page keeps each of the three attached to the population it actually describes, which is the only way the number stays usable.

Three studies, one number
What it actually measuresSourcePopulation
58.5% — share of all US Google searches that end without any clickSparkToro / Datos, 2024All searches, clickstream panel
~58% — reduction in clicks to a top-ranking page when an AI Overview is presentSeparate click-through analysisOnly queries that trigger an AIO
15–25% — estimated organic traffic reduction attributed to zero-click behaviourBain & Company, Dec 2024Survey of consumers, self-reported

Read those rows as a set. The first describes search behaviour overall. The second is a conditional effect on one subset of queries. The third is a modelled business impact from self-reported survey data. Quoting "58%" without saying which one you mean lets a reader assume whichever is most alarming — usually the second, applied to the population of the first.

The practical consequence

"AI Overviews cut your traffic by 58%" is a claim nobody's data supports. It is the conditional click-through effect, applied to a site's whole traffic, using a number borrowed from a different study about all searches. Each ingredient is real. The dish is not.

Three separate studies produced numbers near 58%. They measure completely different things. In circulation they've collapsed into one figure people quote interchangeably — each ingredient real, the dish invented.

Share on X

The same collapse has happened to the zero-click figure, which is quoted for a population it does not cover — traced separately here.

The conversion spread

The second broken chain is more consequential commercially, because it is the number people use to justify budget. Published claims about how much better AI-referred traffic converts range from 4.4× to 23×, with one study reporting 14.2% against 2.8% across 312 B2B firms. These are not small disagreements. They are different orders of magnitude, and they cannot all describe the same phenomenon.

The conversion benchmarks page lists every published claim in that range side by side, with what each one counted.

The likely explanation is not dishonesty. It is that "conversion" is not one thing:

  • Some datasets count any goal completion; others count only revenue events.
  • Some compare AI referrals to all organic; others to non-brand organic only.
  • AI referral volume is tiny — around 1% of traffic in one large panel — so small absolute numbers produce wildly unstable ratios.
  • Sites that show up in AI answers skew toward considered, high-intent purchases in the first place, which inflates the comparison before anyone measures anything.

That last point is the one nobody controls for, and it may account for most of the effect. Until somebody publishes a like-for-like definition, the honest version of this statistic is: AI referral traffic appears to convert better, by an amount nobody has credibly established.

The unverifiable tier

Four claims sit in an awkward middle. They come from named organisations doing real work at genuine scale — a reported 680 million citations is not a number anyone invents casually — but the corpus, the query set and the extraction method are not public.

This is not an accusation. It is a structural observation, and the same one that motivates the AI Citation Index: a company whose product is proprietary citation data cannot open that data without dismantling its own business. The incentive runs the wrong way, permanently. So these figures should be quoted the way you would quote a well-sourced press claim — attributed, dated, and flagged as unverified. Not laundered into a bullet point that reads like established fact.

Why bad statistics spread, and who benefits

The mechanism is structural rather than a matter of a few careless writers. A number with a caveat attached — "roughly," "in one non-public analysis," "for a specific subset of queries" — is less shareable than the same number stated flatly. Every hop has a mild incentive to drop a qualifier, because the flatter version reads more authoritatively. Repeat that across three or four hops and a hedged, population-specific finding becomes an unqualified, universal-sounding fact — not through any single act of dishonesty, but through many small simplifications, each individually reasonable.

ActorWhat they gain from an unqualified numberWhat would change their behaviour
Original researcherUsually nothing extra — most disclose their method fully at the sourceAlready doing the right thing; not the problem
Trade publication summarising the studyA punchier headline, more shares, less reader drop-off from a caveatEditorial norms that reward accurate hedging over false confidence
Content-farm aggregatorA "100+ statistics" page ranks and monetises regardless of whether each figure holds upReaders and links rewarding verification over volume
A practitioner repeating it in a deck or postSounds more authoritative without the extra research time to check itA visible verification norm, like the one this site is arguing for

The first row is not the villain. In every traceable claim in the ledger, the original research was disclosed properly — the degradation happens downstream, at the aggregation and repetition stages, exactly where a reader has the least visibility into what changed.

One propagation vector is specific to this moment. A statistic that appears across enough pages has a real chance of being absorbed into a language model's training data, then repeated on request. It carries the same false confidence as the pages it learned from, now delivered as if it were the model's own synthesis. That is a fifth hop, and the "source" a curious reader is shown may simply restate the same unsourced chain. The implication is unchanged, with higher stakes: trace before you trust, whether the number came from a blog, a slide deck, or a chat window.

What good provenance looks like

What the clean chains had in common
PracticeWhy it makes a claim checkable
Sample size stated in the headline"1.9M citations" tells you immediately what weight the finding bears
Collection window datedThese products change monthly; an undated figure decays silently
Population defined"US searches", "queries that trigger an AIO" — prevents the 58% collision
Method described well enough to disagree withThe test of real disclosure is whether a critic could attack it
Updated figures published as updatesAhrefs re-ran and reported 76% → 38% rather than quietly editing

That last row deserves particular credit. Publishing a revision that undercuts your own earlier headline is the single strongest signal of good faith available to a research publisher, and it costs real attention to do.

This site's own publishing rule

This audit is not commentary. It is the operating rule for every other page here. Any statistic quoted on this site carries the same three-tier grade used in the ledger. A number graded Broken chain is either qualified explicitly in the text or left out, never smoothed into a confident bullet point. Where this site's own research produces a figure — as with the Citation Index — the underlying data is published, so someone else could in principle run this same audit against it. Where every other statistic here sits on this scale is recorded in the master AI SEO statistics index, and a first-party example of the disclosure standard argued for here is the llms.txt log test.

Next step

Apply the five checks to one figure you are currently quoting, then see how this site grades the rest: the studies index carries every method and dataset published here. Each re-run of this audit ships through the newsletter.

Limitations

An audit about provenance has an obligation to be explicit about its own.

  • Twelve claims is a small sample, and it was not randomly drawn. Claims were selected for prominence by judgement. A different selection would produce different proportions, and the 5/4/3 split should be read as a pattern, not a population estimate.
  • Tracing was done through public web search. A chain that looks broken may have a primary source behind a paywall, in a PDF, or on a page that has since moved.
  • "Broken" is a statement about the chain, not the claim. An untraceable statistic may still be perfectly accurate. This audit measures checkability only.
  • It is a snapshot. Pages get edited and studies get published. Grades are true as of the first week of September 2026 and will be re-run twice a year.
  • This page is itself a source that could be mis-cited. If you quote the 5/4/3 split, quote the sample size with it.

What the next edition adds

This is v1 of a recurring audit. The next edition is scheduled alongside the first full run of the Citation Index, and adds three things. It expands the traced-claim set from twelve to at least thirty. It adds a fourth grade tier separating "traceable but superseded" from "traceable and current", because a correctly sourced figure from 2024 can still mislead when quoted as though it describes 2027. And it begins tracking how long a corrected figure, like Ahrefs' 76%-to-38% update, takes to displace the outdated version in circulation. That lag, once measured, will itself be a statistic worth tracing.

How to cite this
Namdev, R. (2026). Where AI SEO statistics actually come from (v1). Retrieved from https://ritiknamdev.com/blog/where-ai-seo-statistics-come-from

Published under CC BY 4.0 — reuse freely with attribution.

FAQ

Frequently asked questions

Is this saying AI SEO statistics are fake?
No — and that would be the lazy version of this piece. Five of the twelve claims traced cleanly to a named study with a disclosed method and sample. The problem is not fabrication. It is that traceable and untraceable numbers circulate side by side, in the same sentence, with no way for a reader to tell them apart.
Why does it matter if a statistic cannot be traced?
Because you cannot know what it measures. Two of the numbers in this audit sound identical and describe completely different things — one is a share of all searches, the other a change in click-through on one kind of result. Quoting either without the method attached means you may be making a claim you did not intend and cannot defend.
Are vendor statistics inherently untrustworthy?
No. Some of the best research in this field comes from commercial tools, and Ahrefs in particular publishes with unusual rigour. The distinction that matters is not commercial versus independent — it is whether the methodology is disclosed well enough that someone could disagree with it.
Isn't tracing twelve claims a small sample to draw conclusions from?
Yes, and the piece says so explicitly in the limitations section. Twelve is enough to demonstrate that the problem is real and to establish a repeatable method — it is not enough to claim a precise, generalizable percentage for the whole field. The next edition expands the set specifically to address this.
Why grade in three tiers instead of a simple pass/fail?
Because the honest picture has three distinct shapes, not two. A claim that names a real organisation but withholds its corpus (partial) is a genuinely different situation from one with literally no traceable origin (broken) — collapsing them into one "fail" bucket would erase a distinction that matters for how much weight to put on the number.
Does a "traceable" grade mean the finding is definitely true?
No — traceable means the method is disclosed well enough to evaluate, not that the finding is beyond dispute. A well-disclosed study can still have a small sample, a narrow population, or a design flaw a critic could reasonably raise. Traceability is the precondition for that conversation happening at all, not a guarantee it goes one way.
What if I published one of the broken-chain claims?
Most people repeating these numbers did so in good faith, which is exactly how a broken chain propagates. If you have a primary source for anything graded broken here, send it and the grade changes — with credit. Re-runs happen twice a year, and corrections are logged rather than silently edited.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.