Reference · Updated quarterly

The GEO Tactic Evidence Scoreboard

Every tactic the industry recommends for getting cited by AI search, graded by the actual evidence behind it — not by how often it's repeated.

Most 'GEO ranking factors' lists present every tactic with equal confidence. This scoreboard doesn't. Some tactics have a controlled study behind them. Most don't. Both kinds get published here — with the difference visible.

Ritik Namdev Ritik Namdev ·Published September 2026 ·14 tactics graded ·13 min read
The short version

Exactly one rigorous, peer-reviewed, causal study of GEO tactics exists in public — the 2024 Princeton paper. Nearly everything else recommended in "how to rank in ChatGPT" content is a plausible hypothesis repeated at a volume that makes it feel settled. This page separates the two.

Why a scoreboard

Search "AI SEO ranking factors" and you'll find dozens of nearly identical lists. Add statistics, use schema, get backlinks, publish original research, keep content fresh, write clear headings. They're presented with uniform confidence, as though each recommendation carries the same evidentiary weight. It doesn't.

Some of these tactics were tested in a controlled benchmark and shown to move a measurable outcome. Others are inferred from what correlates with citation in observational data. Most are simply repeated because everyone else is repeating them. A practitioner deciding where to spend limited time deserves to know which category a given tactic falls into.

Why this distinction matters practically, not just intellectually

Consider two agencies advising the same client. One presents all fourteen tactics as equally important "ranking factors" and recommends the client implement all of them at once. The other, using this scoreboard, tells the client something different. Four tactics, citing sources, adding statistics and quotations, avoiding keyword stuffing, have controlled evidence behind them and should be prioritized. Several more, freshness, author bios, internal linking, are reasonable low-cost bets with no proof either way. One, schema markup, is actively contested and shouldn't be sold as a guaranteed win.

The second agency's client makes better resourcing decisions, sets more realistic expectations, and isn't blindsided if a "guaranteed ranking factor" turns out to do nothing. It was never presented as guaranteed in the first place.

How to read the grades

Fact

Fact — verifiable against a primary, peer-reviewed or fully disclosed source.

Evidence

Evidence — a study exists with a disclosed method and sample, but it's observational, single-vendor, or not independently replicated.

Hypothesis

Hypothesis — widely recommended, mechanistically plausible, but no study located that actually tests it.

Open question

Open — actively contested, with evidence pointing in different directions.

The full scoreboard

14 commonly recommended GEO tactics, graded September 2026
TacticGradeBest evidence located
Add direct quotationsFact Princeton GEO study: +41% visibility lift, 10,000-query benchmark
Add statistics with sourcesFact Princeton GEO study: +30–40% lift
Cite external sourcesFact Princeton GEO study: +30% lift
Avoid keyword stuffingFact Princeton GEO study: underperforms doing nothing
Publish original research/dataEvidence Consistent with Princeton's statistics finding; widely observed correlationally
Rank well on Google (for AI Overviews specifically)Evidence Ahrefs: ~38% of AIO citations rank top-10 (down from 76%) — moderate, weakening correlation
Get cited/mentioned across the web (not just linked)Evidence Vendor-reported correlation between brand mentions and AI Overview visibility — not independently verified
Use structured data / schema markupOpen question Ahrefs observational test found no clear lift; live-fetch tests found schema often ignored — see §7
Keep content fresh / updatedHypothesis Widely asserted; no controlled test located — RCT registered
Add author bios / credentialsHypothesis No controlled test located — RCT registered
Build backlinksHypothesis Reported weaker predictor than brand mentions in available correlational data; no causal test
Improve internal linkingHypothesis No study located
Improve page speedHypothesis No study located; mechanism is unclear for cached/retrieved content — null test registered
Deploy llms.txtEvidence First-party 90-day log test on this site — see linked study; effect on citation, not just crawl, remains open
29%

of the 14 tracked tactics (4 of 14) carry Fact-grade evidence — a controlled, peer-reviewed or fully disclosed test behind them.

This scoreboard

The one causal result

The Princeton-led paper, Aggarwal et al., presented at KDD 2024, remains the field's only peer-reviewed, controlled test of specific content tactics against a measurable visibility outcome. It ran across a 10,000-query benchmark.

Visibility lift by tactic, Princeton GEO study (2024)
Adding quotations
+41%
Adding statistics
+30–40%
Citing sources
+30%
Keyword stuffing
worse than nothing
Source: Aggarwal et al., KDD 2024, arXiv:2311.09735. Effect sizes are from the original benchmark and predate current-generation AI Mode, ChatGPT Search, and Claude web search — treat as directionally reliable, not current-precision.

The keyword-stuffing result is the one worth dwelling on: it didn't just fail to help, it underperformed doing nothing. That's a genuinely useful negative finding, and it's the kind of result this scoreboard exists to preserve rather than let quietly drop out of circulation as the study ages.

Keyword stuffing didn't just fail to improve GEO citation rate in the one controlled study that's tested it — it underperformed doing nothing at all.

Share on X

How the Princeton study actually worked

This single study carries so much evidentiary weight for the field that it's worth describing its method in more detail than a single citation typically allows. The researchers constructed a benchmark of roughly 10,000 real-world queries across multiple topic domains. They then tested a set of content-level interventions: adding quotations, adding statistics, citing sources, restructuring for readability, and, as a deliberate negative control, keyword stuffing, against a generative-answer visibility metric of their own design.

Each intervention was applied to otherwise-comparable content, and the resulting visibility lift measured relative to an unmodified baseline. This is what makes the study's grade of Fact defensible on this scoreboard: a disclosed sample, a disclosed method, peer review, and a built-in negative control that could have shown no effect, or a positive one, but instead showed a clear underperformance. That's exactly the kind of result that's hard to produce by accident or motivated reasoning.

The schema exception

Schema markup is graded Open question rather than settled in either direction, because the actual evidence conflicts:

Evidence

Ahrefs tracked 1,885 pages that added schema markup and found AI citations "barely moved" — observational, but a real, disclosed test.

Evidence

Separate live-fetch tests across five major AI systems found that none of them used information present only in JSON-LD when fetching a page directly — every system extracted visible HTML only.

Hypothesis

Counter-argument: schema may still assist indexing or retrieval stages that happen before the live fetch — this hasn't been ruled out, only the direct-fetch mechanism has been tested against.

This is exactly the kind of question a randomized controlled test could resolve, and it's registered as one of the pre-registered studies in the dedicated schema RCT, part of the AI Citation Index's research program.

What actually moves the needle, on current evidence

If you can only act on the tactics graded Fact or Evidence : cite your sources, include direct quotations and statistics rather than paraphrased claims, publish original data where you can, and don't rely on schema alone to do work that content structure and citation earn. Everything else on the list is a reasonable bet, not a proven lever — which is a different thing to tell a client than "this is a ranking factor."

What this scoreboard argues against, explicitly

Two specific bad habits this page is designed to push back on. First: presenting a Hypothesis-grade tactic to a client or in published content as though it carries Fact-grade certainty. That's the single most common inflation in GEO advice, and the one this scoreboard exists specifically to correct. Second: dismissing a Hypothesis-grade tactic as worthless because it lacks a controlled study. Untested isn't the same as disproven. Several of these tactics remain reasonable to implement for reasons entirely separate from AI citation. Fresh content and clear internal linking both plausibly help classic SEO and user experience, regardless of any GEO effect.

How this differs from classic SEO ranking-factor lists

Classic SEO has decades of large-scale correlation studies, and occasional controlled tests, behind its own ranking-factor discourse. That discourse is itself imperfect, but it rests on a much larger evidence base than GEO currently has. A useful contrast: Google's own ranking system incorporates hundreds of signals refined over two decades of continuous testing at massive scale, observed indirectly through correlation studies and semi-official guidance.

GEO, by comparison, has one controlled academic study and a handful of vendor observational reports. It's a field still in its first few years, where the honest answer to "does X affect AI citation" is far more often "we don't know yet" than current content in the space tends to admit.

How this updates

Re-graded quarterly, or immediately when a new controlled study is published. A tactic's grade moving (up or down) will be logged with the date and the study that caused the change — the same discipline as every other reference asset on this site.

Limitations

  • The Princeton study is aging relative to how fast these platforms change — see the caveat in §2's FAQ and §6's detail.
  • "No study located" is not the same as "no effect exists." Several Hypothesis-grade tactics may well work; they simply haven't been tested in anything we could find.
  • Grades reflect public evidence only. Vendors likely hold internal data that could shift several of these grades if published.
  • This list is not exhaustive. It covers the fourteen most commonly recommended tactics encountered during this site's research, not every GEO tactic ever proposed.
How to cite this
Namdev, R. (2026). The GEO Tactic Evidence Scoreboard (v3). Retrieved from https://ritiknamdev.com/blog/geo-tactic-evidence-scoreboard

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

For the study these grades draw most heavily on, see the Princeton GEO paper referenced below. For a worked example of a tactic tested to a null result, see the planned schema RCT registered in the Citation Index.

FAQ

Frequently asked questions

Where does the evidence grade come from?
Each tactic is graded by the strongest evidence located for it, tiered the same way as every claim on this site: Fact, Evidence, Hypothesis, or Open. A tactic graded Hypothesis is widely recommended, but has no controlled study behind it that we could locate.
Is the Princeton study still relevant given how fast these platforms change?
Its exact effect sizes are from 2023-24 testing, and predate Google AI Mode, ChatGPT Search's current form, and Claude's web search tool entirely. The mechanism it identified, that citable, well-evidenced language performs better, is plausible on priors to still hold. But the specific percentages should be treated as dated until re-tested. That's exactly what our planned GEO replication study will do.
Why does schema get its own section instead of one row?
Because it's the single most contested tactic in the field, with real experiments pointing in different directions. It deserves the nuance a table cell can't hold.
What would move a tactic from Hypothesis to Evidence?
A published study with a disclosed method and sample, testing that specific tactic's effect on citation, even if observational. Fact requires a primary source. Evidence requires disclosed methodology. Neither requires the finding to be causal.
Why does the Princeton study get treated as "Fact" grade rather than just "Evidence"?
Because it's peer-reviewed, discloses its full methodology, and was conducted at a scale, 10,000 queries, that supports the specific effect sizes reported. That's the highest evidentiary bar this site's grading system recognizes, distinct from a single-vendor observational report.
Should a practitioner ignore every tactic graded Hypothesis?
No. A Hypothesis grade means untested, not disproven. Several of these tactics, fresh content, internal linking, are mechanistically plausible and low-cost to implement regardless of AI-citation effects, since they may help classic SEO or user experience independent of the GEO question.
Does this scoreboard cover every possible GEO tactic?
It covers the fourteen most commonly recommended ones encountered across GEO content reviewed for this site. Newer or more niche tactics get added as they become common enough to warrant tracking. See the changelog discipline described in §11.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.