Exactly one rigorous, peer-reviewed, causal study of GEO tactics exists in public — the 2024 Princeton paper. Nearly everything else recommended in "how to rank in ChatGPT" content is a plausible hypothesis repeated at a volume that makes it feel settled. This page separates the two.
Why a scoreboard
Search "AI SEO ranking factors" and you'll find dozens of nearly identical lists. Add statistics, use schema, get backlinks, publish original research, keep content fresh, write clear headings. They're presented with uniform confidence, as though each recommendation carries the same evidentiary weight. It doesn't.
Some of these tactics were tested in a controlled benchmark and shown to move a measurable outcome. Others are inferred from what correlates with citation in observational data. Most are simply repeated because everyone else is repeating them. A practitioner deciding where to spend limited time deserves to know which category a given tactic falls into.
Why this distinction matters practically, not just intellectually
Consider two agencies advising the same client. One presents all fourteen tactics as equally important "ranking factors" and recommends the client implement all of them at once. The other, using this scoreboard, tells the client something different. Four tactics, citing sources, adding statistics and quotations, avoiding keyword stuffing, have controlled evidence behind them and should be prioritized. Several more, freshness, author bios, internal linking, are reasonable low-cost bets with no proof either way. One, schema markup, is actively contested and shouldn't be sold as a guaranteed win.
The second agency's client makes better resourcing decisions, sets more realistic expectations, and isn't blindsided if a "guaranteed ranking factor" turns out to do nothing. It was never presented as guaranteed in the first place.
How to read the grades
Fact — verifiable against a primary, peer-reviewed or fully disclosed source.
Evidence — a study exists with a disclosed method and sample, but it's observational, single-vendor, or not independently replicated.
Hypothesis — widely recommended, mechanistically plausible, but no study located that actually tests it.
Open — actively contested, with evidence pointing in different directions.
The full scoreboard
| Tactic | Grade | Best evidence located |
|---|---|---|
| Add direct quotations | Fact | Princeton GEO study: +41% visibility lift, 10,000-query benchmark |
| Add statistics with sources | Fact | Princeton GEO study: +30–40% lift |
| Cite external sources | Fact | Princeton GEO study: +30% lift |
| Avoid keyword stuffing | Fact | Princeton GEO study: underperforms doing nothing |
| Publish original research/data | Evidence | Consistent with Princeton's statistics finding; widely observed correlationally |
| Rank well on Google (for AI Overviews specifically) | Evidence | Ahrefs: ~38% of AIO citations rank top-10 (down from 76%) — moderate, weakening correlation |
| Get cited/mentioned across the web (not just linked) | Evidence | Vendor-reported correlation between brand mentions and AI Overview visibility — not independently verified |
| Use structured data / schema markup | Open question | Ahrefs observational test found no clear lift; live-fetch tests found schema often ignored — see §7 |
| Keep content fresh / updated | Hypothesis | Widely asserted; no controlled test located — RCT registered |
| Add author bios / credentials | Hypothesis | No controlled test located — RCT registered |
| Build backlinks | Hypothesis | Reported weaker predictor than brand mentions in available correlational data; no causal test |
| Improve internal linking | Hypothesis | No study located |
| Improve page speed | Hypothesis | No study located; mechanism is unclear for cached/retrieved content — null test registered |
| Deploy llms.txt | Evidence | First-party 90-day log test on this site — see linked study; effect on citation, not just crawl, remains open |
of the 14 tracked tactics (4 of 14) carry Fact-grade evidence — a controlled, peer-reviewed or fully disclosed test behind them.
The one causal result
The Princeton-led paper, Aggarwal et al., presented at KDD 2024, remains the field's only peer-reviewed, controlled test of specific content tactics against a measurable visibility outcome. It ran across a 10,000-query benchmark.
The keyword-stuffing result is the one worth dwelling on: it didn't just fail to help, it underperformed doing nothing. That's a genuinely useful negative finding, and it's the kind of result this scoreboard exists to preserve rather than let quietly drop out of circulation as the study ages.
Keyword stuffing didn't just fail to improve GEO citation rate in the one controlled study that's tested it — it underperformed doing nothing at all.
Share on XHow the Princeton study actually worked
This single study carries so much evidentiary weight for the field that it's worth describing its method in more detail than a single citation typically allows. The researchers constructed a benchmark of roughly 10,000 real-world queries across multiple topic domains. They then tested a set of content-level interventions: adding quotations, adding statistics, citing sources, restructuring for readability, and, as a deliberate negative control, keyword stuffing, against a generative-answer visibility metric of their own design.
Each intervention was applied to otherwise-comparable content, and the resulting visibility lift measured relative to an unmodified baseline. This is what makes the study's grade of Fact defensible on this scoreboard: a disclosed sample, a disclosed method, peer review, and a built-in negative control that could have shown no effect, or a positive one, but instead showed a clear underperformance. That's exactly the kind of result that's hard to produce by accident or motivated reasoning.
The schema exception
Schema markup is graded Open question rather than settled in either direction, because the actual evidence conflicts:
Ahrefs tracked 1,885 pages that added schema markup and found AI citations "barely moved" — observational, but a real, disclosed test.
Separate live-fetch tests across five major AI systems found that none of them used information present only in JSON-LD when fetching a page directly — every system extracted visible HTML only.
Counter-argument: schema may still assist indexing or retrieval stages that happen before the live fetch — this hasn't been ruled out, only the direct-fetch mechanism has been tested against.
This is exactly the kind of question a randomized controlled test could resolve, and it's registered as one of the pre-registered studies in the dedicated schema RCT, part of the AI Citation Index's research program.
What actually moves the needle, on current evidence
If you can only act on the tactics graded Fact or Evidence : cite your sources, include direct quotations and statistics rather than paraphrased claims, publish original data where you can, and don't rely on schema alone to do work that content structure and citation earn. Everything else on the list is a reasonable bet, not a proven lever — which is a different thing to tell a client than "this is a ranking factor."
What this scoreboard argues against, explicitly
Two specific bad habits this page is designed to push back on. First: presenting a Hypothesis-grade tactic to a client or in published content as though it carries Fact-grade certainty. That's the single most common inflation in GEO advice, and the one this scoreboard exists specifically to correct. Second: dismissing a Hypothesis-grade tactic as worthless because it lacks a controlled study. Untested isn't the same as disproven. Several of these tactics remain reasonable to implement for reasons entirely separate from AI citation. Fresh content and clear internal linking both plausibly help classic SEO and user experience, regardless of any GEO effect.
How this differs from classic SEO ranking-factor lists
Classic SEO has decades of large-scale correlation studies, and occasional controlled tests, behind its own ranking-factor discourse. That discourse is itself imperfect, but it rests on a much larger evidence base than GEO currently has. A useful contrast: Google's own ranking system incorporates hundreds of signals refined over two decades of continuous testing at massive scale, observed indirectly through correlation studies and semi-official guidance.
GEO, by comparison, has one controlled academic study and a handful of vendor observational reports. It's a field still in its first few years, where the honest answer to "does X affect AI citation" is far more often "we don't know yet" than current content in the space tends to admit.
How this updates
Re-graded quarterly, or immediately when a new controlled study is published. A tactic's grade moving (up or down) will be logged with the date and the study that caused the change — the same discipline as every other reference asset on this site.
Limitations
- The Princeton study is aging relative to how fast these platforms change — see the caveat in §2's FAQ and §6's detail.
- "No study located" is not the same as "no effect exists." Several Hypothesis-grade tactics may well work; they simply haven't been tested in anything we could find.
- Grades reflect public evidence only. Vendors likely hold internal data that could shift several of these grades if published.
- This list is not exhaustive. It covers the fourteen most commonly recommended tactics encountered during this site's research, not every GEO tactic ever proposed.
Namdev, R. (2026). The GEO Tactic Evidence Scoreboard (v3). Retrieved from https://ritiknamdev.com/blog/geo-tactic-evidence-scoreboard Published under CC BY 4.0 — reuse freely with attribution.
For the study these grades draw most heavily on, see the Princeton GEO paper referenced below. For a worked example of a tactic tested to a null result, see the planned schema RCT registered in the Citation Index.