GEO's causal evidence base is one peer-reviewed study: a roughly 10,000-query benchmark run by Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi, tested in 2023–24 and presented at KDD 2024. It found the largest visibility lift — around 41% — from adding direct quotations. Everything else recommended as a GEO tactic rests on correlation, on vendor-reported figures with no public method, or on nothing published at all.
- The +41% visibility lift from adding direct quotations is measured against the researchers' own visibility metric, on their ~10,000-query benchmark, tested 2023–24. It is not a measured lift on any live engine today.
- The study includes a negative control that came back negative — keyword stuffing performed worse than making no change. That is the detail that makes it credible: the measurement could detect a failure.
- 4 of the 14 commonly recommended GEO tactics carry fact-grade evidence, and all four come from that single study. The other ten are recommended daily and untested.
- Two tactics have published null results: llms.txt showed no clear citation effect across ~300,000 domains, and observational work on schema markup found little.
- The Princeton testing predates Google AI Mode, ChatGPT Search in its current form, and Claude web search. Nobody has re-run it. Treat the direction as probably valid and the percentages as dated.
visibility lift from adding direct quotations — the strongest single tactic, measured on the study's own visibility metric across a ~10,000-query benchmark.
commonly recommended GEO tactics with fact-grade evidence behind them. All four trace to the same study.
What this page covers, and what it does not
This page is narrow on purpose. It grades the figures that describe GEO tactics — things a publisher can do to a page — plus the lineage of the one controlled study those figures descend from. It is one scope-limited slice of the master statistics index, where every figure on this site is recorded with its verification status.
It does not grade per-engine citation figures, ranking-overlap percentages, market share, crawler economics or conversion benchmarks. Those live on the AI SEO statistics master index, which grades all sixteen categories on this site and links to the engine-specific pages. If you arrived looking for "how much does Google ranking predict AI citation," that is the page you want, not this one.
The GEO-tactic figures, graded
The grade in the third column is not a judgement of whether the claim is true. It records how far the number can be traced toward a source someone outside the reporting organisation could inspect. A partial grade on a figure you were about to quote is not a reason to drop it — it is a reason to quote it with its qualifier attached.
| Tactic figure | Where to read more | Grade |
|---|---|---|
| Quotations, statistics and cited sources lift visibility 30–41% (~10,000-query benchmark, tested 2023–24) | Tactic scoreboard | Traceable |
| Keyword stuffing performs worse than no change (negative control, same benchmark) | Tactic scoreboard | Traceable |
| llms.txt shows no clear effect on AI citations (~300k domains) | Does llms.txt work | Traceable (a rare published null) |
| Schema markup shows no clear citation lift | Schema study | Traceable (observational precedent) |
| Brand mentions correlate more strongly than backlinks | Brand mentions vs. backlinks | Partial |
| Crawl-to-citation latency | Latency study | Partial (first-party, small sample) |
| Content freshness lifts citation | Freshness study | Partial |
| GEO/AEO/LLMO terminology usage | Terminology tracker | Broken chain (no disclosed-method dataset found) |
Traceable means you can follow the figure to a named, checkable source: a peer-reviewed paper, a disclosed method, a dataset someone else could inspect. Partial means a real, named source exists but the full method or dataset is not public. Broken chain means the trail dead-ends, with the figure repeating across secondary coverage and no disclosed method underneath it. The same three-tier scheme runs on every statistics page here.
What the Princeton study actually measured
Four rows in that table trace to a single source, so it deserves more than a citation. The full paper built a benchmark of roughly 10,000 real-world queries across multiple topic domains, applied isolated content changes to otherwise-comparable material, and measured the resulting visibility against an unmodified baseline. Adding direct quotations produced the largest lift. Adding statistics and citing sources followed. Keyword stuffing, included deliberately as a negative control, performed worse than making no change.
Three features make it unusually strong for this field. It is peer-reviewed, which almost nothing else here is. It discloses its method and sample size. And the negative control actually came back negative, which demonstrates the measurement was capable of detecting a failure rather than finding an effect everywhere it looked.
Two caveats travel with it. The testing dates from 2023–24, predating Google AI Mode, ChatGPT Search in its current form, and Claude's web search tool. The term itself has since become a named discipline with a life well beyond the paper. And the visibility metric was the researchers' own construction rather than an industry-standard measure. That was defensible, since none existed. It is also why the effect sizes are specific to that metric rather than transferable to any engine's live behaviour.
The practical read: trust the direction strongly, treat the percentages as dated, and note that nobody has re-run it. In a field this commercially active, the absence of a replication attempt on its only controlled study is itself worth noticing. What GEO means covers the term and the paper's origin in more depth.
What is recommended versus what has been tested
Content specificity is the one lever with controlled evidence. Quotations, statistics, cited sources — the pattern that original research tends to win on, broadly what correlational factor studies and Google's own guidance point at.
Three widely recommended tactics have weaker evidence than their advocates suggest. Schema markup, where the observational work found little. Freshness, despite confident advice to the contrary. And author bios and E-E-A-T signals. All three are associated with citation in observational data; none has been shown to cause it.
Almost nothing here is causal. One controlled study, two product generations old, underpins the entire causal layer of this discipline. Everything else observes association. That is the honest summary of the evidence base, and it should make anyone quoting these numbers — this site included — more careful rather than less.
The field measures presence, never persistence. Every figure here is a snapshot. Nothing published describes whether an effect lasts, or how long it takes to arrive — the gap our zero-to-cited log study was built to start closing.
GEO's entire causal evidence base is one peer-reviewed study, tested in 2023-24, on the researchers' own visibility metric. Four of fourteen recommended tactics trace to it. The other ten trace to nothing.
Share on XWhy GEO tactic figures are unusually hard to verify
The engines are closed systems that do not publish their retrieval or ranking logic. The companies best placed to measure citation behaviour at scale are often the visibility vendors selling that data as a product, which gives them a direct reason not to fully disclose their method. The field is new, so most cited studies are recent. And the engines keep changing, fast enough that a number measured six months ago may no longer describe current behaviour — the reason our citation half-life study exists at all.
Not every figure ages at the same rate, and this matters more than it sounds. Mechanism findings — that specific, sourced, quotable content gets cited more than vague content — are claims about how retrieval works. They are the part of the Princeton result most likely to still hold.
Negative findings age best of all. Keyword stuffing performing worse than doing nothing is unlikely to reverse, and the llms.txt nulls are shaping up the same way: an Ahrefs study across ~300k domains and independent adoption tracking both point at no detectable citation effect. Negative results age better than positive ones and drop out of circulation faster, which is exactly backwards.
How a good GEO number becomes a bad one
- 1 A qualified finding In a sample of X queries, on engine Y, during window Z, we observed A%.
- 2 A trade article summarises The date window drops first. It makes the sentence long and the article is about now.
- 3 A roundup bullets it The engine qualifier goes. The claim silently widens to every AI engine.
- 4 A fourth writer cites the roundup The link now points at an article, not data. The chain is broken.
- 5 It becomes common knowledge Years old, one engine, quoted about all of them, and never checked again.
The most common failure here is not fabrication. It is degradation: an accurate figure losing the qualifiers that made it accurate, one repetition at a time. Nobody lies at any step — each writer shortens slightly for readability. The cumulative effect is a confident, widely repeated claim that no longer resembles the measurement underneath it. This is the specific mechanism the grade column exists to interrupt, and why a grade travels with every figure on this site rather than sitting in a methodology note nobody reads. The full hop-by-hop tracing method is in the provenance audit.
Red flags in a GEO statistics roundup
Apply these to any GEO statistics page, this one included. The last one does most of the work: a roundup that never flags a single figure as uncertain, in a field this contested, is a roundup where nobody tried to verify anything. Not every roundup fails these checks — Ahrefs' own collection at least names its sources — but most in the genre do.
Two habits follow directly. Never state a partial-graded figure to one decimal place: if the method is not public, the precision is not defensible, and "roughly half" is more honest than "47.9%" even though the second sounds more authoritative. And do not build a strategy on a single figure. Where several independent sources point the same direction, that direction is worth acting on. Where one vendor reports one number, it is worth knowing and not worth reorganising a content programme around.
The GEO measurements that should exist and do not
| Missing measurement | Why it matters | Is it hard, or just undone? |
|---|---|---|
| A replication of the Princeton study | It is the entire causal layer of the discipline, and it is two product generations old | Undone. The method is published. |
| Per-engine effect sizes | Engines disagree on what they cite; one aggregate figure may describe none of them | Undone, and cheap at small scale |
| Effect sizes for the other ten tactics | A dozen tactics get recommended daily with no measured effect at all | Undone |
| Query-type segmentation | Definitional, comparison and how-to queries almost certainly behave differently | Undone |
| Durability of any effect | Every figure is a snapshot; strategy assumes persistence nobody has tested | Hard — it needs a longitudinal panel |
| Independent check of any vendor figure | No tactic-correlation number has been verified by a party with no product to sell | Hard — the corpora are the product |
The most valuable of these is the first. Re-running the same interventions against current engines would either confirm the field's foundation or reveal that it has shifted, and the method is already published. The most instructive is the last. Classic SEO's evidence base improved because multiple independent parties ran large studies, published methods, and disagreed with each other publicly. Disagreement between disclosed methods is productive; agreement between undisclosed ones tells you nothing. Brand-level measurement is stuck in exactly that state — the mentions-to-visibility gap is described everywhere and quantified nowhere.
Where our own data will replace these
As the Citation Index releases data, the partial-graded rows above get rebuilt on first-party numbers and re-graded. Everything published lands in the studies index, including results we would rather not have found, which go in the null results registry.
Those releases will carry the raw rows, so anyone can recompute rather than trusting the summary. That is the specific difference intended between a first-party figure here and the vendor-reported figures it replaces — not that ours will be more accurate, which nobody can promise in advance, but that ours will be checkable. If a first-party measurement contradicts a vendor figure in this table, both get published side by side with methods stated. Quietly swapping in a more convenient number is exactly the practice this page exists to document.
The one GEO tactic with a published null result is easy to test on your own site: generate an llms.txt and read what the evidence actually says about it before you assume it does anything. New grades ship via the newsletter.
Verification status
Grades on this page follow the same hop-by-hop tracing method described above, applied to its own figures as well as everyone else's. Each figure is traced to a primary source wherever possible, and graded partial or broken where the trail stops short. Who does the grading, and on what basis, is set out with the method.
Namdev, R. (2026). GEO statistics: the tactic figures and the Princeton lineage (v1). Retrieved from https://ritiknamdev.com/blog/geo-statistics Published under CC BY 4.0 — reuse freely with attribution.
AI SEO statistics is the master index for every statistics category on this site, including the per-engine citation figures deliberately excluded here. AEO statistics covers the separate question of what has been published under the answer-engine label.