About 40% of the field's most-quoted statistics trace cleanly to a named study with a disclosed method. The rest either name a source that withholds its methodology, or dead-end entirely. The failure mode is not fabrication — it is that checkable and uncheckable numbers circulate in the same sentence with nothing to distinguish them.
The five provenance checks
These are the checks used to grade every figure in this audit, and they are the grading scheme applied to every statistic published anywhere on this site. Each takes about two minutes. Together they catch nearly everything in the ledger below.
- Click the citation, not the summary. Follow the actual link, one hop at a time, until you reach a primary source, a circular loop, or a page that cites nothing. If there is no link at all, stop there — you already have your answer.
- Find the population. A percentage without a denominator is not a statistic. "US searches" and "queries that trigger an AI Overview" are different worlds, and conflating them is how the 58% collision below happens.
- Find the date and the collection window. These products change monthly. An AI-search figure older than roughly six months may describe a product that no longer exists in that form, and an undated figure decays silently.
- Check the method is disclosed well enough to disagree with. Sample size, technique, and how the measurement was taken. The real test of disclosure is whether a critic could attack it — not whether an organisation's name appears.
- Ask what would falsify it. If no possible observation would contradict the claim, it is a position rather than a finding, however many decimal places it carries.
Grades follow directly from those checks. Traceable means the chain reaches a named study with sample, date and method disclosed. Partial means it reaches a named organisation whose underlying method or corpus is withheld. Broken chain means no primary source is reachable, or the sources contradict each other irreconcilably. And when a number passes all five, cite the primary source rather than the page you found it on — every extra hop is one more place the method can fall out.
A completed provenance audit. Each figure below was traced by hand to its earliest locatable source; the method and the full result set are on this page.
- Which of the field's most-quoted statistics trace cleanly to a named study with a disclosed method, and which do not.
- Where each broken chain terminates — the specific point at which provenance stops.
- The five checks used to grade a figure, published so anyone can apply them.
- Whether the untraceable figures are wrong — an untraceable number is unverifiable, not disproven.
- Whether the traced sample is representative of every figure in circulation; it covers the most-quoted ones, not all of them.
An unchecked statistic is a liability in a specific way. When you quote a figure in a client deck or a strategy doc, you are implicitly vouching for what it measures. It may instead measure something adjacent — a different population, a different time window, a different definition of the same word. The recommendation built on top of it is then wrong in a way nobody can see until it fails.
- A budget reallocated on "AI traffic converts 23× better" — one of three unreconciled vendor claims.
- A crawler blocked on a crawl-to-referral ratio drawn from a different vertical entirely.
- A journalist repeating "AI Overviews cut clicks by 58%", a figure this audit found conflates two separate studies.
None of these require bad faith. They require only that a plausible-sounding number circulated widely enough, for long enough, that checking it stopped feeling necessary.
So this audit asks one question of twelve claims: can a reader get from the statistic to the thing that produced it? Not "is it true" — that is a much harder question. Just: is the path there open.
An unchecked statistic is a liability. When you quote a figure, you vouch for what it measures — and if it measures something adjacent, the recommendation built on it is wrong in a way nobody can see until it fails.
Share on XHow we traced twelve claims
The method is deliberately dull, because the value is in doing it consistently rather than cleverly.
- 01 Collect the claim As phrased on a page that repeats it, verbatim
- 02 Follow the link Whatever the page cites, one hop at a time
- 03 Repeat to the end Until a primary source, a loop, or a dead end
- 04 Check the method Is sample, date and technique disclosed?
- 05 Grade and record Traceable / partial / broken, with the chain
What we found
- Traceable to primary 5
- Partial — source named, method withheld 4
- Broken chain 3
The headline is less damning and more interesting than "SEO statistics are made up." Nearly half trace cleanly. The problem is the other half — and specifically that a reader has no way to tell which half a given number belongs to, because both are presented in the same confident register.
of the twelve traced claims graded fully traceable to a named study with disclosed method and sample.
Chain length is the practical signal. One or two hops is normal and healthy — a trade publication summarising a study is doing its job. Past three hops, the original method has almost always been stripped away, and what survives is a number and a vibe.
claims traced cleanly to a named study with a disclosed sample, method and date.
claims where no primary source was reachable, or where sources contradict irreconcilably.
distinct studies found circulating under the same '58%' figure, each measuring something different.
The full ledger
Every claim, its grade, and where the chain actually ends. This is the part worth bookmarking.
| Claim as commonly quoted | Chain ends at | Method disclosed? | Grade |
|---|---|---|---|
| "76% of AI Overview citations come from top-10 pages" — later "down to 38%" | Ahrefs, Jul 2025 (1.9M citations) and Mar 2026 (863K SERPs) | Yes — sample and window stated | Traceable |
| "Only 12% of AI-cited URLs rank in Google's top 10" | Ahrefs, 2026 | Yes | Traceable |
| "58.5% of US searches end without a click" | SparkToro with Datos clickstream, 2024 | Yes — panel and geography stated | Traceable |
| "GEO tactics lift visibility up to 40%" | Aggarwal et al., KDD 2024 — peer-reviewed, 10,000-query benchmark | Yes — fully specified | Traceable |
| "80% of consumers rely on zero-click results in 40%+ of searches" | Bain & Company survey, Dec 2024 | Yes — survey design stated | Traceable |
| "AI crawlers take 2,237 pages per referral" (and similar ratios) | Cloudflare network data — but usually quoted via aggregators, not Radar itself | Partly — the derivation is rarely shown | Partial |
| "ChatGPT's citations are 47.9% Wikipedia" | Profound, from a reported 680M-citation corpus | No — corpus and query set not public | Partial |
| "Perplexity cites Reddit 46.7% of the time" | Same corpus as above | No | Partial |
| "Only 11% of domains are cited by both ChatGPT and Perplexity" | Vendor analysis; query set not published | No | Partial |
| "AI Overviews cut clicks by 58%" | A different study from the 58.5% zero-click figure — routinely conflated with it | Varies by retelling; the conflation destroys the meaning | Broken chain |
| "AI traffic converts 4.4× better" / "23× better" / "14.2% vs 2.8%" | Three separate vendor datasets that cannot be reconciled | Partially, but definitions of "conversion" differ | Broken chain |
| Round-number market-size and adoption forecasts | Content-farm pages citing each other; no primary data located in the loop | No | Broken chain |
A live thirteenth case, not yet in the ledger, shows the same method working on a fresh claim. "Google AI Mode fans a query out into 8–16 sub-queries" is repeated across dozens of SEO articles from 2025 and 2026, almost always without a citation. Google's documentation confirms the fan-out mechanism exists, and a granted patent (US11663201B2) describes eight distinct sub-query types the system may generate. Nowhere in that patent, or in any Google publication located during this research, is a specific per-query count of "8 to 16" stated. Graded Broken chain — plausible, oft-repeated, untraceable to any primary count. Full write-up at the Query Fan-Out Corpus study.
One trace, in full: how "2,237 pages per referral" was checked
Ledgers are abstract. One trace, step by step, exactly as it happened.
- The claim, as encountered: a marketing blog post stated "Anthropic's ClaudeBot crawls 2,237 pages for every visitor it sends back," with no link attached.
- Hop one: a search for the exact phrase surfaced a second aggregator site repeating the same number, this time with a link to a third-party SEO-tools blog.
- Hop two: that SEO-tools blog cited "Cloudflare data, July 2026" in prose, again without a direct link to Cloudflare Radar itself, and rounded the figure slightly differently (2,200-ish) than the first source.
- Hop three: Cloudflare Radar's own public dashboard was checked directly. It exposes bot-traffic classification and category-level breakdowns. The specific per-crawler ratio, as phrased, was not found on Radar as a standalone published figure. A third party most likely derived it from Radar's underlying data rather than quoting a Cloudflare publication.
- Verdict: graded Partial . The underlying network is almost certainly Cloudflare's, and the figure is plausible against the general pattern Cloudflare has described publicly. But nobody in the chain shows the derivation — which time window, which subset of traffic, how "referral" was defined. A reader repeating "2,237:1" is repeating a number nobody has shown their work for.
That is what "partial" looks like in practice: not a fabricated number, but one that has lost its method somewhere between the primary network data and the blog post that made it famous.
The 58% problem
This is the clearest finding in the audit, because it shows how a statistic can be simultaneously accurate and useless. Three separate, legitimate pieces of research produced numbers near 58%. They measure entirely different things. In circulation, they have collapsed into one figure people quote interchangeably. The CTR statistics page keeps each of the three attached to the population it actually describes, which is the only way the number stays usable.
| What it actually measures | Source | Population |
|---|---|---|
| 58.5% — share of all US Google searches that end without any click | SparkToro / Datos, 2024 | All searches, clickstream panel |
| ~58% — reduction in clicks to a top-ranking page when an AI Overview is present | Separate click-through analysis | Only queries that trigger an AIO |
| 15–25% — estimated organic traffic reduction attributed to zero-click behaviour | Bain & Company, Dec 2024 | Survey of consumers, self-reported |
Read those rows as a set. The first describes search behaviour overall. The second is a conditional effect on one subset of queries. The third is a modelled business impact from self-reported survey data. Quoting "58%" without saying which one you mean lets a reader assume whichever is most alarming — usually the second, applied to the population of the first.
"AI Overviews cut your traffic by 58%" is a claim nobody's data supports. It is the conditional click-through effect, applied to a site's whole traffic, using a number borrowed from a different study about all searches. Each ingredient is real. The dish is not.
Three separate studies produced numbers near 58%. They measure completely different things. In circulation they've collapsed into one figure people quote interchangeably — each ingredient real, the dish invented.
Share on XThe same collapse has happened to the zero-click figure, which is quoted for a population it does not cover — traced separately here.
The conversion spread
The second broken chain is more consequential commercially, because it is the number people use to justify budget. Published claims about how much better AI-referred traffic converts range from 4.4× to 23×, with one study reporting 14.2% against 2.8% across 312 B2B firms. These are not small disagreements. They are different orders of magnitude, and they cannot all describe the same phenomenon.
The conversion benchmarks page lists every published claim in that range side by side, with what each one counted.
The likely explanation is not dishonesty. It is that "conversion" is not one thing:
- Some datasets count any goal completion; others count only revenue events.
- Some compare AI referrals to all organic; others to non-brand organic only.
- AI referral volume is tiny — around 1% of traffic in one large panel — so small absolute numbers produce wildly unstable ratios.
- Sites that show up in AI answers skew toward considered, high-intent purchases in the first place, which inflates the comparison before anyone measures anything.
That last point is the one nobody controls for, and it may account for most of the effect. Until somebody publishes a like-for-like definition, the honest version of this statistic is: AI referral traffic appears to convert better, by an amount nobody has credibly established.
The unverifiable tier
Four claims sit in an awkward middle. They come from named organisations doing real work at genuine scale — a reported 680 million citations is not a number anyone invents casually — but the corpus, the query set and the extraction method are not public.
This is not an accusation. It is a structural observation, and the same one that motivates the AI Citation Index: a company whose product is proprietary citation data cannot open that data without dismantling its own business. The incentive runs the wrong way, permanently. So these figures should be quoted the way you would quote a well-sourced press claim — attributed, dated, and flagged as unverified. Not laundered into a bullet point that reads like established fact.
Why bad statistics spread, and who benefits
The mechanism is structural rather than a matter of a few careless writers. A number with a caveat attached — "roughly," "in one non-public analysis," "for a specific subset of queries" — is less shareable than the same number stated flatly. Every hop has a mild incentive to drop a qualifier, because the flatter version reads more authoritatively. Repeat that across three or four hops and a hedged, population-specific finding becomes an unqualified, universal-sounding fact — not through any single act of dishonesty, but through many small simplifications, each individually reasonable.
| Actor | What they gain from an unqualified number | What would change their behaviour |
|---|---|---|
| Original researcher | Usually nothing extra — most disclose their method fully at the source | Already doing the right thing; not the problem |
| Trade publication summarising the study | A punchier headline, more shares, less reader drop-off from a caveat | Editorial norms that reward accurate hedging over false confidence |
| Content-farm aggregator | A "100+ statistics" page ranks and monetises regardless of whether each figure holds up | Readers and links rewarding verification over volume |
| A practitioner repeating it in a deck or post | Sounds more authoritative without the extra research time to check it | A visible verification norm, like the one this site is arguing for |
The first row is not the villain. In every traceable claim in the ledger, the original research was disclosed properly — the degradation happens downstream, at the aggregation and repetition stages, exactly where a reader has the least visibility into what changed.
One propagation vector is specific to this moment. A statistic that appears across enough pages has a real chance of being absorbed into a language model's training data, then repeated on request. It carries the same false confidence as the pages it learned from, now delivered as if it were the model's own synthesis. That is a fifth hop, and the "source" a curious reader is shown may simply restate the same unsourced chain. The implication is unchanged, with higher stakes: trace before you trust, whether the number came from a blog, a slide deck, or a chat window.
What good provenance looks like
| Practice | Why it makes a claim checkable |
|---|---|
| Sample size stated in the headline | "1.9M citations" tells you immediately what weight the finding bears |
| Collection window dated | These products change monthly; an undated figure decays silently |
| Population defined | "US searches", "queries that trigger an AIO" — prevents the 58% collision |
| Method described well enough to disagree with | The test of real disclosure is whether a critic could attack it |
| Updated figures published as updates | Ahrefs re-ran and reported 76% → 38% rather than quietly editing |
That last row deserves particular credit. Publishing a revision that undercuts your own earlier headline is the single strongest signal of good faith available to a research publisher, and it costs real attention to do.
This site's own publishing rule
This audit is not commentary. It is the operating rule for every other page here. Any statistic quoted on this site carries the same three-tier grade used in the ledger. A number graded Broken chain is either qualified explicitly in the text or left out, never smoothed into a confident bullet point. Where this site's own research produces a figure — as with the Citation Index — the underlying data is published, so someone else could in principle run this same audit against it. Where every other statistic here sits on this scale is recorded in the master AI SEO statistics index, and a first-party example of the disclosure standard argued for here is the llms.txt log test.
Apply the five checks to one figure you are currently quoting, then see how this site grades the rest: the studies index carries every method and dataset published here. Each re-run of this audit ships through the newsletter.
Limitations
An audit about provenance has an obligation to be explicit about its own.
- Twelve claims is a small sample, and it was not randomly drawn. Claims were selected for prominence by judgement. A different selection would produce different proportions, and the 5/4/3 split should be read as a pattern, not a population estimate.
- Tracing was done through public web search. A chain that looks broken may have a primary source behind a paywall, in a PDF, or on a page that has since moved.
- "Broken" is a statement about the chain, not the claim. An untraceable statistic may still be perfectly accurate. This audit measures checkability only.
- It is a snapshot. Pages get edited and studies get published. Grades are true as of the first week of September 2026 and will be re-run twice a year.
- This page is itself a source that could be mis-cited. If you quote the 5/4/3 split, quote the sample size with it.
What the next edition adds
This is v1 of a recurring audit. The next edition is scheduled alongside the first full run of the Citation Index, and adds three things. It expands the traced-claim set from twelve to at least thirty. It adds a fourth grade tier separating "traceable but superseded" from "traceable and current", because a correctly sourced figure from 2024 can still mislead when quoted as though it describes 2027. And it begins tracking how long a corrected figure, like Ahrefs' 76%-to-38% update, takes to displace the outdated version in circulation. That lag, once measured, will itself be a statistic worth tracing.
Namdev, R. (2026). Where AI SEO statistics actually come from (v1). Retrieved from https://ritiknamdev.com/blog/where-ai-seo-statistics-come-from Published under CC BY 4.0 — reuse freely with attribution.