This is a hub, not a duplicate. Every statistic below links to its own dedicated page with full context, method, and a verification grade. The table lets you scan the field fast and jump to what matters.
Headline numbers
Two figures anchor most of what is known. One comes from the field's only controlled study. The other describes how weakly classic ranking now predicts AI citation, and varies widely by engine.
visibility lift from adding direct quotations, the strongest single tactic in the one peer-reviewed GEO study.
range of AI-citation-to-Google-ranking overlap, varying substantially by engine.
GEO statistics, verified
Each row links to a dedicated page carrying the full method, sample and caveats. The grade in the third column is not a judgement of whether the claim is true. It records how far the number can be traced toward a source someone outside the reporting organisation could actually inspect. A partial grade on a figure you were about to quote is not a reason to drop it. It is a reason to quote it with its qualifier attached, which is a habit almost nobody in this field practises.
| Statistic | Where to read more | Grade |
|---|---|---|
| Quotations/statistics lift GEO visibility 30–41% | Tactic scoreboard | Traceable |
| Schema markup shows no clear citation lift | Schema RCT | Traceable (observational precedent) |
| ChatGPT/Perplexity domain overlap ~11% | Most-cited domains | Partial |
| Brand mentions correlate more strongly than backlinks | Brand mentions vs. backlinks | Partial |
| GEO/AEO/LLMO terminology usage | Terminology tracker | Broken chain (no disclosed-method dataset found) |
| AI Overview citation-to-ranking overlap fell ~76% → ~38% YoY | AI Overview statistics | Traceable |
| Cross-platform citation concordance ~11% (2-engine, vendor-reported) | Concordance study | Partial |
The Princeton figures, in detail
Four rows in that table trace to a single source. Given how much weight it carries, it deserves more than a citation.
The study built a benchmark of roughly 10,000 real-world queries across multiple topic domains. It then applied isolated content changes to otherwise-comparable material and measured the resulting visibility against an unmodified baseline. Adding direct quotations produced the largest lift. Adding statistics and citing sources followed. Keyword stuffing, included deliberately as a negative control, performed worse than making no change at all.
Three features make it unusually strong for this field. It is peer-reviewed, which almost nothing else here is. It discloses its method and sample size. And the negative control actually came back negative, which demonstrates the measurement was capable of detecting a failure rather than finding an effect everywhere it looked.
Two caveats travel with it. The testing dates from 2023-24, predating Google AI Mode, ChatGPT Search in its current form, and Claude's web search tool. And the visibility metric was the researchers' own construction rather than an industry-standard measure, which is defensible given none existed, and still means the effect sizes are specific to that metric.
The practical read: trust the direction strongly, treat the percentages as dated, and note that nobody has re-run it. In a field this commercially active, the absence of a replication attempt on its only controlled study is itself worth noticing.
The ranking-overlap figures, in detail
The second-strongest family of figures here concerns how much AI citation overlaps with classic Google ranking. These come from Ahrefs, with disclosed samples, which is why they grade traceable.
The headline: overlap varies enormously by engine. AI Overviews sit far higher than ChatGPT, Gemini and Copilot, which get grouped together in the source study at a much lower figure. Perplexity is reported separately as the outlier, tracking Google rank more closely than the other non-Google engines.
That spread is the actual finding, and it gets flattened constantly in coverage. A single "AI citations only overlap with Google top-10 about 12% of the time" claim describes one grouping of engines. Applied to AI Overviews it is simply wrong, by a wide margin.
The trend figure is stronger still, and rarer. The AI Overview overlap fell year over year, measured by the same organisation using a comparable method both times. That combination is what permits a trend claim at all. Most "X changed to Y" statements in this field pair two numbers produced by different methods, which makes the change uninterpretable.
What these figures do not tell you: whether ranking causes citation. High overlap is consistent with ranking driving citation, with both being driven by a third factor like content quality, or with the engines drawing on a shared underlying index. The correlation is measured. The mechanism is not.
The source-skew figures, in detail
The most quotable numbers on this page are also its weakest. Every source-skew figure, which domain types dominate which engine's citations, comes from vendor-reported analysis of a corpus that is not independently public.
They are graded partial rather than broken for a specific reason. A named vendor stands behind them, the corpus size is stated, and the figures are internally consistent with each other and with observable engine behaviour. That is meaningfully better than an unsourced number. It is meaningfully worse than something anyone can check.
The structural problem is unavoidable and worth naming. The organisations capable of analysing citations at this scale built that capability to sell a product. Publishing the corpus would give away the product. So the method stays partially described, and no outside party can reproduce the analysis. This is not misconduct. It is an incentive structure, and it explains most of why this page contains so many partial grades.
How to use them anyway: treat the ordering as more reliable than the magnitudes. That ChatGPT skews encyclopedic and Perplexity skews community is consistent across sources and matches observable behaviour. That the specific figure is 47.9% rather than 45% or 52% rests entirely on one unpublished corpus.
Why GEO statistics are unusually hard to verify
GEO sits in an unusual spot. The engines themselves are closed systems. They don't publish their retrieval or ranking logic. And the companies best placed to measure citation behavior at scale are often the same visibility vendors selling that data as a product. That gives them a direct reason not to fully disclose their method.
Add two more problems on top. The field is new. Most cited studies are from the past two years. And the engines themselves keep changing, fast enough that a number measured six months ago may no longer describe today's behavior. Together, that makes GEO statistics harder to verify than most other marketing topics. It's exactly why the grading system on this page exists, instead of presenting every number at face value.
How to read the grading column
The grade is the most important column in that table, and it is the one most likely to be skimmed past. It answers a different question than the statistic does. The figure tells you what someone measured. The grade tells you how much weight the measurement can bear.
Traceable means you can follow the figure to a named, checkable source. A peer-reviewed paper. A disclosed method. A dataset someone else could inspect.
Partial means a real source exists, but the full method or dataset isn't public. A named vendor stands behind the number, but nobody outside that company can fully audit how it was made. Broken chain means the trail dead-ends. No disclosed-method source exists at all, only the figure repeating across secondary coverage. This three-tier system is the same across every statistics page on this site.
Most 'GEO statistics 2026' pages present every number with the same confidence. This one grades each figure traceable, partial, or broken — because that difference is the whole point.
Share on XWhat the numbers add up to
Read individually, these figures are a collection of disconnected facts. Read together, they describe a coherent and fairly specific picture. It is worth stating that picture explicitly, because no single row in the table conveys it.
Classic ranking is a weakening predictor, not a dead one. The overlap figures fell sharply on Google's own AI surface and sit low across most non-Google engines. Ranking still helps. It is no longer close to sufficient, and for several engines it was never the main driver.
The engines are genuinely different systems. Different source skews, different ranking overlaps, low reported concordance between them. That combination means "AI visibility" as a single target is a category error, and it is the most consistently supported conclusion available from this data.
Content specificity is the one lever with controlled evidence. Quotations, statistics, cited sources. Not schema, not freshness, not author bios, all of which are recommended more confidently than the evidence supports. The gap between what is recommended and what is demonstrated is wide.
The field measures presence, never persistence. Every figure here is a snapshot. Nothing published describes whether an effect lasts. That is a striking hole given how much strategic advice implicitly assumes durability.
Almost nothing is causal. One controlled study, two product generations old, underpins the entire causal layer of this discipline. Everything else observes association. That is the honest summary of the evidence base, and it should make anyone quoting these numbers, including this site, more careful rather than less.
Where our own data will replace these
Several rows above rely on vendor-reported figures. No independent, reproducible dataset exists yet for them. As the Citation Index releases its own data, those rows will get rebuilt on first-party numbers and re-graded to match.
Three rows are the priority, chosen because they are both weakly sourced and heavily quoted. The source-skew percentages, currently resting entirely on one unpublished corpus. The cross-platform concordance figure, currently covering two engines from one vendor. And the citation-density comparisons between engines, which are described qualitatively everywhere and measured nowhere.
When those land, they will be published with the raw rows attached, so anyone can recompute them rather than trusting the summary. That is the specific difference intended between a first-party figure here and the vendor-reported figures it replaces. Not that ours will be more accurate, which nobody can promise in advance, but that ours will be checkable.
If a first-party measurement contradicts a vendor figure currently in this table, both get published side by side with the methods stated. Quietly replacing a number with a more convenient one, and not mentioning the disagreement, is exactly the practice this page exists to document.
How a good number becomes a bad one
The most common failure in this field is not fabrication. It is degradation: an accurate figure losing the qualifiers that made it accurate, one repetition at a time. The process is predictable enough to describe step by step.
Step one: an original study publishes a qualified finding. Something like "in a sample of X queries, on engine Y, during window Z, we observed A%." Fully qualified, entirely defensible.
Step two: a trade article summarises it. The window drops first, because it makes the sentence long and the article is about the present. Now it reads "on engine Y, A% of citations..."
Step three: a roundup includes it as a bullet. Bullets are short. The engine qualifier goes. Now it reads "A% of AI citations..." and describes something far broader than what was measured.
Step four: a fourth writer cites the roundup. The link now points to the roundup, not the study. Anyone trying to verify hits an article, not data. The chain is broken and the figure looks just as authoritative as it did at step one.
Step five: it becomes common knowledge. The number appears in enough places that checking feels unnecessary. At this point it may be years old, describe one engine, and be quoted about all of them.
Nobody lied at any step. Each writer shortened slightly for readability. The cumulative effect is a confident, widely-repeated claim that no longer resembles the measurement underneath it. This is the specific mechanism the grading column exists to interrupt, and it is why a grade travels with every figure on this site rather than sitting in a methodology note nobody reads.
Red flags in a GEO statistics roundup
Watch for these patterns in any GEO statistics page, this one included. A number with no named source at all. A source link that, once you check it, turns out to be another blog post rather than original data. A suspiciously precise figure, like "73.2%," that implies far more measurement accuracy than the method could actually support.
And one more: a roundup that never once flags a figure as uncertain. Given how contested this whole field is, that's a signal that nobody actually tried to verify anything.
A worked example: spotting a red flag
The checks above are quick to describe and easy to skip. Running one all the way through makes the difference concrete.
Say a page states: "GEO improves AI citation rate by exactly 34.7%." No link, no study named. That single sentence already fails two of the checks above at once. There's no named source, and the number is oddly precise for a claim with no visible method behind it.
Compare that to how this page states its own top figure: "+41% visibility lift, from adding direct quotations, per the one peer-reviewed GEO study (Princeton, KDD 2024)." Same shape of claim. But this one names the study, names the specific tactic, and links out so you can check it yourself.
The statistics that should exist and do not
A statistics page is defined partly by its holes. These are the numbers a practitioner would most want, which nobody has produced.
A replication of the Princeton study. The single most valuable missing measurement. Its effect sizes are the closest thing this field has to causal evidence, and they are two product generations old. Re-running the same interventions against current engines would either confirm the field's foundation or reveal that it has shifted.
Per-tactic effect sizes beyond the original three. The study tested a handful of content changes. A dozen more tactics get recommended daily with no measured effect at all. Even one additional controlled test would meaningfully change the evidence landscape.
Effect sizes broken down by engine. Every existing GEO effect size is an aggregate. Given how much engines appear to disagree on what they cite, a tactic that lifts visibility on one may do nothing on another. No published figure separates them.
Query-type segmentation. Do these tactics work equally for definitional queries, comparisons, and how-to questions? Almost certainly not. No figure exists that splits by query type.
Anything measuring durability. Every GEO statistic describes a moment. None describe whether an effect persists. A tactic that produces a two-week lift and one that produces a lasting change are indistinguishable in all current data.
Independent verification of any vendor figure. Not a new measurement, just a second one. No source-skew number on this page has ever been checked by a party without a commercial interest in the answer.
Using these numbers without misusing them
Practical rules for anyone about to quote something from this table.
Carry the grade with the number. Not as a footnote. In the same sentence. "Vendor-reported from a non-public corpus" costs six words and preempts the only serious objection anyone will raise.
Never state a partial-graded figure to one decimal place. If the method is not public, the precision is not defensible. "Roughly half" is more honest than "47.9%" when the underlying corpus cannot be examined, even though the second sounds more authoritative.
Name the engine. Almost every figure here describes one specific engine or one grouping of engines. Dropping that qualifier is the single most common way an accurate GEO statistic becomes a false one.
Do not build a strategy on a single figure. Where multiple independent sources point the same direction, that direction is worth acting on. Where one vendor reports one number, that number is worth knowing and not worth reorganising a content programme around.
Check the linked page before publishing. Grades move. A figure that was traceable last quarter may have been superseded by a re-measurement using a different method.
How fast these figures go stale
Every number here has a shelf life, and they are not the same length. Sorting by how quickly a figure becomes unreliable is more useful than sorting by how solid it looked when published.
Fastest to expire: source-skew percentages. These describe what an engine currently prefers to cite. Engines change retrieval behaviour frequently, and at least one documented case shows a major platform's reliance on a single source swinging dramatically within weeks. Treat any skew figure older than a few months as directional at best.
Moderately durable: ranking-overlap figures. These describe a structural relationship between two systems rather than a momentary preference. The observed year-over-year decline shows they do move, but they move on a timescale of quarters rather than weeks.
Most durable: mechanism findings. That specific, sourced, quotable content gets cited more than vague content is a claim about how retrieval works, not about a particular engine's current configuration. Even if the Princeton effect sizes are dated, the mechanism they identified is the part most likely to still hold.
Effectively permanent: the negative findings. Keyword stuffing performing worse than doing nothing is unlikely to reverse. Negative results tend to age better than positive ones, and they also tend to drop out of circulation faster, which is exactly backwards.
The practical implication: when you find a GEO statistic without a date attached, that omission tells you something. Figures in this field are not timeless, and the ones quoted most confidently are frequently the ones that expire fastest.
How this compares to a mature evidence base
A useful way to calibrate how much to trust anything on this page: compare the state of GEO evidence against a field that has had time to mature.
Classic SEO ranking research is imperfect and frequently criticised, and it rests on decades of large-scale correlation studies, repeated measurement by multiple independent parties, and occasional controlled tests. Practitioners argue about what the data means. They rarely argue about whether data exists.
GEO has one controlled study and a set of vendor reports. That is not a criticism of anyone working in the field. It is what a discipline looks like two or three years in. Classic SEO was in a comparable state at a comparable age, with confident advice substantially outrunning the evidence for it.
The instructive part is what changed. Classic SEO's evidence base improved because multiple independent parties ran large studies, published methods, and disagreed with each other publicly. Disagreement between disclosed methods is productive. Agreement between undisclosed ones tells you nothing.
That suggests what would most improve this field, and it is not more advice. It is a second party measuring the same thing a first party measured, with both methods public, so the results can be compared. Currently almost every figure in GEO exists in exactly one version, produced once, by one organisation.
Verification status
This page's own grades follow the same procedure it recommends readers apply to everyone else's numbers, including its own.
Grades follow the same method as the provenance audit. Each figure is traced hop by hop to a primary source, wherever that's possible. It gets graded partial or broken where the trail stops short.
Namdev, R. (2026). GEO statistics (v1). Retrieved from https://ritiknamdev.com/blog/geo-statistics Published under CC BY 4.0 — reuse freely with attribution.
See the GEO tactic evidence scoreboard for tactic-level grading, and where AI SEO statistics actually come from for the provenance methodology used throughout this page.