About 40% of the field's most-quoted statistics trace cleanly to a named study with a disclosed method. The rest either name a source that withholds its methodology, or dead-end entirely. The failure mode is not fabrication — it is that checkable and uncheckable numbers circulate in the same sentence with nothing to distinguish them.
Why trace anything
Every field that runs on numbers eventually has to audit its own. SEO has not done this for AI search, and the reason is understandable: the numbers are useful, they sound authoritative, and checking one properly takes half an hour that nobody is paying for.
But an unchecked statistic is a liability in a specific way. When you quote a figure in a client deck or a strategy doc, you are implicitly vouching for what it measures. If it turns out to measure something adjacent — a different population, a different time window, a different definition of the same word — the recommendation built on top of it may be wrong in a way nobody can see until it fails.
So this audit asks one question of twelve claims: can a reader get from the statistic to the thing that produced it? Not "is it true" — that is a much harder question. Just: is the path there open.
An unchecked statistic is a liability. When you quote a figure, you vouch for what it measures — and if it measures something adjacent, the recommendation built on it is wrong in a way nobody can see until it fails.
Share on XWhy this matters practically, not just academically
It would be easy to treat this as a pedantic exercise — "well, actually, that statistic isn't quite right" is not a popular sentence in a client meeting. But consider three concrete situations where an untraced statistic causes a real, costly decision:
- A budget conversation. A marketing lead cites "AI traffic converts 23× better" to justify reallocating spend away from paid search. If that figure is one of three unreconciled vendor claims — see the conversion spread below — the actual number for that business could be closer to break-even, and the reallocation loses money.
- A technical decision. A site owner blocks an AI crawler based on a crawl-to-referral ratio quoted without its source explained, not realising the figure came from a different industry vertical entirely, with a very different traffic profile than their own site.
- A public claim. A journalist or analyst repeats "AI Overviews cut clicks by 58%" in a published piece, unaware — as this audit found — that the figure conflates two entirely separate studies. The correction, if it comes, arrives long after the original claim has spread.
None of these outcomes require anyone to have acted in bad faith. They require only that a plausible-sounding number circulated widely enough, for long enough, that checking it stopped feeling necessary.
How we traced them
The method is deliberately dull, because the value is in doing it consistently rather than cleverly.
- 01 Collect the claim As phrased on a page that repeats it, verbatim
- 02 Follow the link Whatever the page cites, one hop at a time
- 03 Repeat to the end Until a primary source, a loop, or a dead end
- 04 Check the method Is sample, date and technique disclosed?
- 05 Grade and record Traceable / partial / broken, with the chain
What we found
- Traceable to primary 5
- Partial — source named, method withheld 4
- Broken chain 3
The headline is less damning and more interesting than "SEO statistics are made up." Nearly half trace cleanly. The problem is the other half — and specifically that a reader has no way to tell which half a given number belongs to, because both are presented in the same confident register.
of the twelve traced claims graded fully traceable to a named study with disclosed method and sample.
Chain length is the practical signal. One or two hops is normal and healthy — a trade publication summarising a study is doing its job. Past three hops, the original method has almost always been stripped away, and what survives is a number and a vibe.
claims traced cleanly to a named study with a disclosed sample, method and date.
claims where no primary source was reachable, or where sources contradict irreconcilably.
distinct studies found circulating under the same '58%' figure, each measuring something different.
The full ledger
Every claim, its grade, and where the chain actually ends. This is the part worth bookmarking.
| Claim as commonly quoted | Chain ends at | Method disclosed? | Grade |
|---|---|---|---|
| "76% of AI Overview citations come from top-10 pages" — later "down to 38%" | Ahrefs, Jul 2025 (1.9M citations) and Mar 2026 (863K SERPs) | Yes — sample and window stated | Traceable |
| "Only 12% of AI-cited URLs rank in Google's top 10" | Ahrefs, 2026 | Yes | Traceable |
| "58.5% of US searches end without a click" | SparkToro with Datos clickstream, 2024 | Yes — panel and geography stated | Traceable |
| "GEO tactics lift visibility up to 40%" | Aggarwal et al., KDD 2024 — peer-reviewed, 10,000-query benchmark | Yes — fully specified | Traceable |
| "80% of consumers rely on zero-click results in 40%+ of searches" | Bain & Company survey, Dec 2024 | Yes — survey design stated | Traceable |
| "AI crawlers take 2,237 pages per referral" (and similar ratios) | Cloudflare network data — but usually quoted via aggregators, not Radar itself | Partly — the derivation is rarely shown | Partial |
| "ChatGPT's citations are 47.9% Wikipedia" | Profound, from a reported 680M-citation corpus | No — corpus and query set not public | Partial |
| "Perplexity cites Reddit 46.7% of the time" | Same corpus as above | No | Partial |
| "Only 11% of domains are cited by both ChatGPT and Perplexity" | Vendor analysis; query set not published | No | Partial |
| "AI Overviews cut clicks by 58%" | A different study from the 58.5% zero-click figure — routinely conflated with it | Varies by retelling; the conflation destroys the meaning | Broken chain |
| "AI traffic converts 4.4× better" / "23× better" / "14.2% vs 2.8%" | Three separate vendor datasets that cannot be reconciled | Partially, but definitions of "conversion" differ | Broken chain |
| Round-number market-size and adoption forecasts | Content-farm pages citing each other; no primary data located in the loop | No | Broken chain |
One trace, in full: how "2,237 pages per referral" was checked
Ledgers are abstract. It helps to walk through one trace step by step, exactly as it happened, so the method is visible rather than just described.
- The claim, as encountered: a marketing blog post stated "Anthropic's ClaudeBot crawls 2,237 pages for every visitor it sends back," with no link attached.
- Hop one: a search for the exact phrase surfaced a second aggregator site repeating the same number, this time with a link to a third-party SEO-tools blog.
- Hop two: that SEO-tools blog cited "Cloudflare data, July 2026" in prose, again without a direct link to Cloudflare Radar itself, and rounded the figure slightly differently (2,200-ish) than the first source.
- Hop three: Cloudflare Radar's own public dashboard was checked directly. It exposes bot-traffic classification and category-level breakdowns, but the specific per-crawler ratio, as phrased, was not found presented as a standalone published figure on Radar itself — meaning the number likely was derived by a third party from Radar's underlying data rather than quoted verbatim from a Cloudflare publication.
- Verdict: graded Partial . The underlying network almost certainly is Cloudflare's, and the figure is plausible given the general pattern Cloudflare has publicly described, but the exact derivation — which time window, which subset of traffic, how "referral" was defined — is not shown by anyone in the chain. A reader repeating "2,237:1" is repeating a number nobody has shown their work for.
This is what "partial" looks like in practice: not a fabricated number, but one that has lost its method somewhere between the primary network data and the blog post that made it famous.
The incentive map: who benefits from an unqualified number
It helps to separate the actors in this chain by what they actually gain from dropping a qualifier, because "everyone is equally at fault" is both untrue and unhelpful for deciding where to apply scrutiny.
| Actor | What they gain from an unqualified number | What would change their behaviour |
|---|---|---|
| Original researcher | Usually nothing extra — most disclose their method fully at the source | Already doing the right thing; not the problem |
| Trade publication summarising the study | A punchier headline, more shares, less reader drop-off from a caveat | Editorial norms that reward accurate hedging over false confidence |
| Content-farm aggregator | A "100+ statistics" page ranks and monetises regardless of whether each figure holds up | Readers and links rewarding verification over volume |
| A practitioner repeating it in a deck or post | Sounds more authoritative without the extra research time to check it | A visible verification norm, like the one this site is arguing for |
Note that the first row is not the villain of this story. In every traceable claim in the ledger above, the original research was disclosed properly — the degradation happens downstream, at the aggregation and repetition stages, which is exactly where a reader has the least visibility into what changed.
A live example: the query fan-out folklore
A useful test of this whole framework is to apply it to a claim not yet in the main ledger: "Google AI Mode fans a query out into 8–16 sub-queries." This figure is repeated across dozens of SEO articles published in 2025 and 2026, almost always without a citation. Tracing it: Google's own documentation confirms the fan-out mechanism exists and a granted patent (US11663201B2) describes eight distinct sub-query types the system may generate. Nowhere in that patent, or in any Google publication located during this research, is a specific per-query count of "8 to 16" stated. The number appears to have originated as an estimate by an early commentator, then been repeated at increasing confidence with each retelling until it reads as an official figure. Graded Broken chain — plausible, oft-repeated, and untraceable to any primary count. The full write-up, including a proposed corpus that would replace the estimate with real measurement, is at the Query Fan-Out Corpus study.
Why bad statistics spread faster than good ones
It is worth understanding the mechanism, because it explains why this problem is structural rather than a matter of a few careless writers. A number with a caveat attached — "roughly," "in one non-public analysis," "for a specific subset of queries" — is less shareable than the same number stated flatly. Every hop in a citation chain has a mild incentive to drop a qualifier, because the flatter version reads more authoritatively and performs better as a headline or a bullet point. Repeat that pressure across three or four hops and a hedged, population-specific finding turns into an unqualified, universal-sounding fact — not through any single act of dishonesty, but through the cumulative effect of many small simplifications, each individually reasonable.
The 58% problem
This is the single clearest finding in the audit, and it is worth sitting with, because it shows how a statistic can be simultaneously accurate and useless.
Three separate, legitimate pieces of research produced numbers near 58%. They measure entirely different things. In circulation, they have collapsed into one figure that people quote interchangeably.
| What it actually measures | Source | Population |
|---|---|---|
| 58.5% — share of all US Google searches that end without any click | SparkToro / Datos, 2024 | All searches, clickstream panel |
| ~58% — reduction in clicks to a top-ranking page when an AI Overview is present | Separate click-through analysis | Only queries that trigger an AIO |
| 15–25% — estimated organic traffic reduction attributed to zero-click behaviour | Bain & Company, Dec 2024 | Survey of consumers, self-reported |
Read those rows as a set. The first is a description of search behaviour overall. The second is a conditional effect on one subset of queries. The third is a modelled business impact from self-reported survey data. Quoting "58%" without saying which one you mean lets a reader assume whichever is most alarming — usually the second, applied to the population of the first.
"AI Overviews cut your traffic by 58%" is a claim nobody's data supports. It is the conditional click-through effect, applied to a site's whole traffic, using a number borrowed from a different study about all searches. Each ingredient is real. The dish is not.
Three separate studies produced numbers near 58%. They measure completely different things. In circulation they've collapsed into one figure people quote interchangeably — each ingredient real, the dish invented.
Share on XThe conversion spread
The second broken chain is more consequential commercially, because it is the number people use to justify budget.
Published claims about how much better AI-referred traffic converts range from 4.4× to 23×, with one study reporting 14.2% against 2.8% across 312 B2B firms. These are not small disagreements. They are different orders of magnitude, and they cannot all describe the same phenomenon.
The likely explanation is not dishonesty. It is that "conversion" is not one thing:
- Some datasets count any goal completion; others count only revenue events.
- Some compare AI referrals to all organic; others to non-brand organic only.
- AI referral volume is tiny — around 1% of traffic in one large panel — so small absolute numbers produce wildly unstable ratios.
- Sites that show up in AI answers skew toward considered, high-intent purchases in the first place, which inflates the comparison before anyone measures anything.
That last point is the one nobody controls for, and it may account for most of the effect. Until somebody publishes a like-for-like definition, the honest version of this statistic is: AI referral traffic appears to convert better, by an amount nobody has credibly established.
The unverifiable tier
Four claims sit in an awkward middle. They come from named organisations doing real work at genuine scale — a reported 680 million citations is not a number anyone invents casually — but the corpus, the query set and the extraction method are not public.
This is not an accusation. It is a structural observation, and it is the same one that motivates the AI Citation Index: a company whose product is proprietary citation data cannot open that data without dismantling its own business. The incentive runs the wrong way, permanently.
So these figures should be quoted the way you would quote a well-sourced press claim — attributed, dated, and flagged as unverified. Not laundered into a bullet point that reads like established fact.
What good provenance looks like
The traceable half of this audit shares a pattern worth naming, because it is easy to copy and almost nobody does.
| Practice | Why it makes a claim checkable |
|---|---|
| Sample size stated in the headline | "1.9M citations" tells you immediately what weight the finding bears |
| Collection window dated | These products change monthly; an undated figure decays silently |
| Population defined | "US searches", "queries that trigger an AIO" — prevents the 58% collision |
| Method described well enough to disagree with | The test of real disclosure is whether a critic could attack it |
| Updated figures published as updates | Ahrefs re-ran and reported 76% → 38% rather than quietly editing |
That last row deserves particular credit. Publishing a revision that undercuts your own earlier headline is the single strongest signal of good faith available to a research publisher, and it costs real attention to do.
How to use a statistic responsibly
Four checks, roughly two minutes each. They catch nearly everything in this audit.
And when a number passes all four, cite the primary source rather than the page you found it on. Every extra hop you add to the chain is one more place the method can fall out.
A new propagation vector: AI systems repeating unsourced statistics
There is a wrinkle specific to this moment that a provenance audit written five years ago would not have needed to address. Large language models are trained on a snapshot of the web, which means an unsourced statistic that has spread widely enough to appear across many pages has a real chance of being absorbed into a model's training data and then repeated, on request, with the same false confidence as the pages it learned from — but now delivered as if it were the model's own synthesis rather than a copy-pasted number.
This does not make AI systems uniquely culpable; they are reproducing a pattern already present in their training data, not inventing it. But it does raise the stakes of the underlying problem, because a chat interface answering a direct question carries an air of authority that a blog post with a visible byline and publish date does not always convey as strongly. A statistic that has gone through four hops of human aggregation and then been absorbed into a model's weights has, in effect, gained a fifth hop — one where the "source" a curious reader is shown, if any, may simply be a restatement of the same unsourced chain rather than a path back to the primary study.
The practical implication is the same as everywhere else in this piece, just with higher stakes: trace before you trust, regardless of whether the confident-sounding number came from a blog, a slide deck, or a chat window.
How this affects what gets published on this site
This audit is not just a piece of commentary — it is the operating rule for every other page on this site. Any statistic quoted here carries the same three-tier grade used throughout this ledger, and a number graded Broken chain is either qualified explicitly in the text or left out entirely rather than smoothed into a confident bullet point. Where this site's own research produces a figure — as with the Citation Index — the underlying data is published so the same audit could, in principle, be run against it by someone else.
Limitations
An audit about provenance has an obligation to be explicit about its own.
- Twelve claims is a small sample, and it was not randomly drawn. Claims were selected for prominence by judgement. A different selection would produce different proportions, and the 5/4/3 split should be read as a pattern, not a population estimate.
- Tracing was done through public web search. A chain that looks broken may have a primary source behind a paywall, in a PDF, or on a page that has since moved.
- "Broken" is a statement about the chain, not the claim. An untraceable statistic may still be perfectly accurate. This audit measures checkability only.
- It is a snapshot. Pages get edited and studies get published. Grades are true as of the first week of September 2026 and will be re-run twice a year.
- This page is itself a source that could be mis-cited. If you quote the 5/4/3 split, quote the sample size with it.
What the next edition will add
This is v1 of a recurring audit, not a one-time exercise. The next edition, scheduled alongside the first full run of the Citation Index, will expand the traced-claim set from twelve to at least thirty, add a fourth grade tier distinguishing "traceable but superseded" from "traceable and current" — since a correctly-sourced figure from 2024 can still mislead if quoted as though it describes 2027 — and begin tracking how long it takes a corrected figure (like Ahrefs' 76%-to-38% update) to actually displace the outdated version in circulation. That lag, once measured, will itself be a statistic worth tracing.
Namdev, R. (2026). Where AI SEO statistics actually come from (v1). Retrieved from https://ritiknamdev.com/blog/where-ai-seo-statistics-come-from Published under CC BY 4.0 — reuse freely with attribution.
This audit is the reason the AI Citation Index publishes its query set and raw data rather than only its conclusions. For the technical groundwork behind the crawler figures above, see GPTBot vs OAI-SearchBot; for a first-party example of the disclosure standard argued for here, see the 90-day llms.txt log test. For where every other statistic on this site sits on this same grading scale, see the master AI SEO statistics index.