Statistics · GEO tactics only

GEO statistics: the tactic figures and the Princeton lineage

The GEO-tactic statistics that actually have a source — the peer-reviewed Princeton effect sizes, the published nulls, and the vendor-reported figures that sit between them. Narrow by design: per-engine citation figures live on the master index.

Ritik Namdev Ritik Namdev ·Published September 2026 ·Scope: GEO tactics ·11 min read ·Last verified September 2026
The short version

GEO's causal evidence base is one peer-reviewed study: a roughly 10,000-query benchmark run by Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi, tested in 2023–24 and presented at KDD 2024. It found the largest visibility lift — around 41% — from adding direct quotations. Everything else recommended as a GEO tactic rests on correlation, on vendor-reported figures with no public method, or on nothing published at all.

What this page establishes
  • The +41% visibility lift from adding direct quotations is measured against the researchers' own visibility metric, on their ~10,000-query benchmark, tested 2023–24. It is not a measured lift on any live engine today.
  • The study includes a negative control that came back negative — keyword stuffing performed worse than making no change. That is the detail that makes it credible: the measurement could detect a failure.
  • 4 of the 14 commonly recommended GEO tactics carry fact-grade evidence, and all four come from that single study. The other ten are recommended daily and untested.
  • Two tactics have published null results: llms.txt showed no clear citation effect across ~300,000 domains, and observational work on schema markup found little.
  • The Princeton testing predates Google AI Mode, ChatGPT Search in its current form, and Claude web search. Nobody has re-run it. Treat the direction as probably valid and the percentages as dated.
+41%

visibility lift from adding direct quotations — the strongest single tactic, measured on the study's own visibility metric across a ~10,000-query benchmark.

Aggarwal et al., KDD 2024 · Tested 2023–24
4 of 14

commonly recommended GEO tactics with fact-grade evidence behind them. All four trace to the same study.

This site's tactic scoreboard · September 2026

What this page covers, and what it does not

This page is narrow on purpose. It grades the figures that describe GEO tactics — things a publisher can do to a page — plus the lineage of the one controlled study those figures descend from. It is one scope-limited slice of the master statistics index, where every figure on this site is recorded with its verification status.

It does not grade per-engine citation figures, ranking-overlap percentages, market share, crawler economics or conversion benchmarks. Those live on the AI SEO statistics master index, which grades all sixteen categories on this site and links to the engine-specific pages. If you arrived looking for "how much does Google ranking predict AI citation," that is the page you want, not this one.

The GEO-tactic figures, graded

The grade in the third column is not a judgement of whether the claim is true. It records how far the number can be traced toward a source someone outside the reporting organisation could inspect. A partial grade on a figure you were about to quote is not a reason to drop it — it is a reason to quote it with its qualifier attached.

Tactic figureWhere to read moreGrade
Quotations, statistics and cited sources lift visibility 30–41% (~10,000-query benchmark, tested 2023–24)Tactic scoreboardTraceable
Keyword stuffing performs worse than no change (negative control, same benchmark)Tactic scoreboardTraceable
llms.txt shows no clear effect on AI citations (~300k domains)Does llms.txt workTraceable (a rare published null)
Schema markup shows no clear citation liftSchema studyTraceable (observational precedent)
Brand mentions correlate more strongly than backlinksBrand mentions vs. backlinksPartial
Crawl-to-citation latencyLatency studyPartial (first-party, small sample)
Content freshness lifts citationFreshness studyPartial
GEO/AEO/LLMO terminology usageTerminology trackerBroken chain (no disclosed-method dataset found)

Traceable means you can follow the figure to a named, checkable source: a peer-reviewed paper, a disclosed method, a dataset someone else could inspect. Partial means a real, named source exists but the full method or dataset is not public. Broken chain means the trail dead-ends, with the figure repeating across secondary coverage and no disclosed method underneath it. The same three-tier scheme runs on every statistics page here.

What the Princeton study actually measured

Four rows in that table trace to a single source, so it deserves more than a citation. The full paper built a benchmark of roughly 10,000 real-world queries across multiple topic domains, applied isolated content changes to otherwise-comparable material, and measured the resulting visibility against an unmodified baseline. Adding direct quotations produced the largest lift. Adding statistics and citing sources followed. Keyword stuffing, included deliberately as a negative control, performed worse than making no change.

Three features make it unusually strong for this field. It is peer-reviewed, which almost nothing else here is. It discloses its method and sample size. And the negative control actually came back negative, which demonstrates the measurement was capable of detecting a failure rather than finding an effect everywhere it looked.

Two caveats travel with it. The testing dates from 2023–24, predating Google AI Mode, ChatGPT Search in its current form, and Claude's web search tool. The term itself has since become a named discipline with a life well beyond the paper. And the visibility metric was the researchers' own construction rather than an industry-standard measure. That was defensible, since none existed. It is also why the effect sizes are specific to that metric rather than transferable to any engine's live behaviour.

The practical read: trust the direction strongly, treat the percentages as dated, and note that nobody has re-run it. In a field this commercially active, the absence of a replication attempt on its only controlled study is itself worth noticing. What GEO means covers the term and the paper's origin in more depth.

Content specificity is the one lever with controlled evidence. Quotations, statistics, cited sources — the pattern that original research tends to win on, broadly what correlational factor studies and Google's own guidance point at.

Three widely recommended tactics have weaker evidence than their advocates suggest. Schema markup, where the observational work found little. Freshness, despite confident advice to the contrary. And author bios and E-E-A-T signals. All three are associated with citation in observational data; none has been shown to cause it.

Almost nothing here is causal. One controlled study, two product generations old, underpins the entire causal layer of this discipline. Everything else observes association. That is the honest summary of the evidence base, and it should make anyone quoting these numbers — this site included — more careful rather than less.

The field measures presence, never persistence. Every figure here is a snapshot. Nothing published describes whether an effect lasts, or how long it takes to arrive — the gap our zero-to-cited log study was built to start closing.

GEO's entire causal evidence base is one peer-reviewed study, tested in 2023-24, on the researchers' own visibility metric. Four of fourteen recommended tactics trace to it. The other ten trace to nothing.

Share on X

Why GEO tactic figures are unusually hard to verify

The engines are closed systems that do not publish their retrieval or ranking logic. The companies best placed to measure citation behaviour at scale are often the visibility vendors selling that data as a product, which gives them a direct reason not to fully disclose their method. The field is new, so most cited studies are recent. And the engines keep changing, fast enough that a number measured six months ago may no longer describe current behaviour — the reason our citation half-life study exists at all.

Not every figure ages at the same rate, and this matters more than it sounds. Mechanism findings — that specific, sourced, quotable content gets cited more than vague content — are claims about how retrieval works. They are the part of the Princeton result most likely to still hold.

Negative findings age best of all. Keyword stuffing performing worse than doing nothing is unlikely to reverse, and the llms.txt nulls are shaping up the same way: an Ahrefs study across ~300k domains and independent adoption tracking both point at no detectable citation effect. Negative results age better than positive ones and drop out of circulation faster, which is exactly backwards.

How a good GEO number becomes a bad one

Five steps from a qualified finding to a false one, with nobody lying
  1. 1 A qualified finding In a sample of X queries, on engine Y, during window Z, we observed A%.
  2. 2 A trade article summarises The date window drops first. It makes the sentence long and the article is about now.
  3. 3 A roundup bullets it The engine qualifier goes. The claim silently widens to every AI engine.
  4. 4 A fourth writer cites the roundup The link now points at an article, not data. The chain is broken.
  5. 5 It becomes common knowledge Years old, one engine, quoted about all of them, and never checked again.

The most common failure here is not fabrication. It is degradation: an accurate figure losing the qualifiers that made it accurate, one repetition at a time. Nobody lies at any step — each writer shortens slightly for readability. The cumulative effect is a confident, widely repeated claim that no longer resembles the measurement underneath it. This is the specific mechanism the grade column exists to interrupt, and why a grade travels with every figure on this site rather than sitting in a methodology note nobody reads. The full hop-by-hop tracing method is in the provenance audit.

Red flags in a GEO statistics roundup

No named sourceA figure with no study, vendor or dataset attached anywhere on the page.
A link that leads to another blog postFollow it once. If it lands on secondary coverage rather than data, the trail is already broken.
Suspicious precisionA figure like 73.2% implies measurement accuracy the disclosed method cannot support.
No engine namedAlmost every real figure describes one engine or one grouping. A bare AI citations claim has lost its population.
No date on anythingThese figures expire. An undated one cannot be checked against the engine it described.
Nothing flagged as uncertainA roundup where every number is equally confident is a roundup where nothing was verified.

Apply these to any GEO statistics page, this one included. The last one does most of the work: a roundup that never flags a single figure as uncertain, in a field this contested, is a roundup where nobody tried to verify anything. Not every roundup fails these checks — Ahrefs' own collection at least names its sources — but most in the genre do.

Two habits follow directly. Never state a partial-graded figure to one decimal place: if the method is not public, the precision is not defensible, and "roughly half" is more honest than "47.9%" even though the second sounds more authoritative. And do not build a strategy on a single figure. Where several independent sources point the same direction, that direction is worth acting on. Where one vendor reports one number, it is worth knowing and not worth reorganising a content programme around.

The GEO measurements that should exist and do not

Missing measurementWhy it mattersIs it hard, or just undone?
A replication of the Princeton studyIt is the entire causal layer of the discipline, and it is two product generations oldUndone. The method is published.
Per-engine effect sizesEngines disagree on what they cite; one aggregate figure may describe none of themUndone, and cheap at small scale
Effect sizes for the other ten tacticsA dozen tactics get recommended daily with no measured effect at allUndone
Query-type segmentationDefinitional, comparison and how-to queries almost certainly behave differentlyUndone
Durability of any effectEvery figure is a snapshot; strategy assumes persistence nobody has testedHard — it needs a longitudinal panel
Independent check of any vendor figureNo tactic-correlation number has been verified by a party with no product to sellHard — the corpora are the product

The most valuable of these is the first. Re-running the same interventions against current engines would either confirm the field's foundation or reveal that it has shifted, and the method is already published. The most instructive is the last. Classic SEO's evidence base improved because multiple independent parties ran large studies, published methods, and disagreed with each other publicly. Disagreement between disclosed methods is productive; agreement between undisclosed ones tells you nothing. Brand-level measurement is stuck in exactly that state — the mentions-to-visibility gap is described everywhere and quantified nowhere.

Where our own data will replace these

As the Citation Index releases data, the partial-graded rows above get rebuilt on first-party numbers and re-graded. Everything published lands in the studies index, including results we would rather not have found, which go in the null results registry.

Those releases will carry the raw rows, so anyone can recompute rather than trusting the summary. That is the specific difference intended between a first-party figure here and the vendor-reported figures it replaces — not that ours will be more accurate, which nobody can promise in advance, but that ours will be checkable. If a first-party measurement contradicts a vendor figure in this table, both get published side by side with methods stated. Quietly swapping in a more convenient number is exactly the practice this page exists to document.

Next step

The one GEO tactic with a published null result is easy to test on your own site: generate an llms.txt and read what the evidence actually says about it before you assume it does anything. New grades ship via the newsletter.

Verification status

Grades on this page follow the same hop-by-hop tracing method described above, applied to its own figures as well as everyone else's. Each figure is traced to a primary source wherever possible, and graded partial or broken where the trail stops short. Who does the grading, and on what basis, is set out with the method.

How to cite this
Namdev, R. (2026). GEO statistics: the tactic figures and the Princeton lineage (v1). Retrieved from https://ritiknamdev.com/blog/geo-statistics

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

AI SEO statistics is the master index for every statistics category on this site, including the per-engine citation figures deliberately excluded here. AEO statistics covers the separate question of what has been published under the answer-engine label.

FAQ

Frequently asked questions

Which of these figures are the most reliable?
The Princeton GEO study figures are traceable and peer-reviewed. Everything else on this page is either a published null result or a vendor-reported figure graded partial.
Why do so many GEO statistics roundups repeat the exact same numbers?
Most trace back to a small handful of original sources: one peer-reviewed study and a few vendor blog posts. Writers cite them, then re-cite each other, until the original source gets hard to find. The provenance audit page documents this pattern in detail.
Can I still quote a partial-graded figure?
Yes, with its qualifier attached. Say who reported it and note the method is not public. That is one extra clause and it is the difference between a claim that survives scrutiny and one that collapses when someone asks for the source.
Why does the Princeton study carry so much weight on this page?
Because it is the only peer-reviewed, controlled test of GEO tactics that exists. Every other figure here is observational or vendor-reported. That is not a compliment to the study so much as a statement about how thin the rest of the evidence base is.
Are the Princeton effect sizes still accurate?
Treat the mechanism as probably still valid and the exact percentages as dated. The testing predates Google AI Mode, current ChatGPT Search, and Claude web search entirely. Nobody has re-run it, which is itself one of the field's more surprising gaps.
What is the difference between this page and the tactic scoreboard?
This page grades figures. The scoreboard grades tactics. A statistic can be well-sourced while the tactic it describes remains unproven, and a tactic can be plausible while every number attached to it is weak. The two grading systems answer different questions.
Where do the per-engine citation figures live?
Not here. This page is deliberately narrow: GEO tactics and the Princeton lineage. Per-engine citation and ranking-overlap figures are graded on the AI SEO statistics master index and on the individual engine pages it links to.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.