Exactly one rigorous, peer-reviewed, causal study of GEO tactics exists in public — the 2024 Princeton paper. Nearly everything else recommended in "how to rank in ChatGPT" content is a plausible hypothesis repeated at a volume that makes it feel settled. This page separates the two.
Why a scoreboard
Search "AI SEO ranking factors" and you'll find dozens of nearly identical lists. Add statistics, use schema, get backlinks, publish original research, keep content fresh, write clear headings. They're presented with uniform confidence, as though each recommendation carries the same evidentiary weight. It doesn't.
Some of these tactics were tested in a controlled benchmark and shown to move a measurable outcome. Others are inferred from what correlates with citation in observational data. Most are simply repeated because everyone else is repeating them. A practitioner deciding where to spend limited time deserves to know which category a given tactic falls into. The same audit applied to the field's numbers rather than its tactics is the provenance audit, and the vocabulary used here is defined in the glossary.
Why this distinction matters practically, not just intellectually
Consider two agencies advising the same client. One presents all fourteen tactics as equally important "ranking factors" and recommends the client implement all of them at once. The other, using this scoreboard, tells the client something different. Four tactics, citing sources, adding statistics and quotations, avoiding keyword stuffing, have controlled evidence behind them and should be prioritized. Several more, freshness, author bios, internal linking, are reasonable low-cost bets with no proof either way. One, schema markup, is actively contested and shouldn't be sold as a guaranteed win.
The second agency's client makes better resourcing decisions, sets more realistic expectations, and isn't blindsided if a "guaranteed ranking factor" turns out to do nothing. It was never presented as guaranteed in the first place. Reporting those grades honestly is what the measurement standard is for, and what the null results registry preserves.
How to read the grades
Fact — verifiable against a primary, peer-reviewed or fully disclosed source.
Evidence — a study exists with a disclosed method and sample, but it's observational, single-vendor, or not independently replicated.
Hypothesis — widely recommended, mechanistically plausible, but no study located that actually tests it.
Open — actively contested, with evidence pointing in different directions.
The full scoreboard
| Tactic | Grade | Last regraded | Best evidence located |
|---|---|---|---|
| Add direct quotations | Fact | Sep 2026 | Princeton GEO study: +41% visibility lift, 10,000-query benchmark |
| Add statistics with sources | Fact | Sep 2026 | Princeton GEO study: +30–40% lift |
| Cite external sources | Fact | Sep 2026 | Princeton GEO study: +30% lift |
| Avoid keyword stuffing | Fact | Sep 2026 | Princeton GEO study: underperforms doing nothing |
| Publish original research/data | Evidence | Sep 2026 | Consistent with Princeton's statistics finding; widely observed correlationally |
| Rank well on Google (for AI Overviews specifically) | Evidence | Sep 2026 | Ahrefs: ~38% of AIO citations rank top-10 (down from 76%) — moderate, weakening correlation |
| Get cited/mentioned across the web (not just linked) | Evidence | Sep 2026 | Vendor-reported correlation between brand mentions and AI Overview visibility — not independently verified |
| Use structured data / schema markup | Open question | Sep 2026 | Ahrefs observational test found no clear lift; live-fetch tests found schema often ignored — see §7 |
| Keep content fresh / updated | Hypothesis | Sep 2026 | Widely asserted; no controlled test located — RCT registered |
| Add author bios / credentials | Hypothesis | Sep 2026 | No controlled test located — RCT registered |
| Build backlinks | Hypothesis | Sep 2026 | Reported weaker predictor than brand mentions in available correlational data; no causal test |
| Improve internal linking | Hypothesis | Sep 2026 | No study located |
| Improve page speed | Hypothesis | Sep 2026 | No study located; mechanism is unclear for cached/retrieved content — null test registered |
| Deploy llms.txt | Evidence | Sep 2026 | First-party 90-day log test on this site — see linked study; effect on citation, not just crawl, remains open |
- Fact — controlled, peer-reviewed or fully disclosed 4
- Evidence — disclosed method, but observational 4
- Hypothesis — widely recommended, no study located 5
- Open — evidence points in different directions 1
Row by row, the deeper treatments are: schema markup, content freshness, author and E-E-A-T signals and llms.txt. The correlational evidence behind the middle rows comes from brand-mention analyses, ranking-factor correlations and, on freshness, vendor coverage with no isolated test behind it.
of the 14 tracked tactics (4 of 14) carry Fact-grade evidence — a controlled, peer-reviewed or fully disclosed test behind them.
The one causal result
The Princeton-led paper, Aggarwal et al., presented at KDD 2024, remains the field's only peer-reviewed, controlled test of specific content tactics against a measurable visibility outcome. It ran across a 10,000-query benchmark, and the full paper publishes the interventions and the negative control. It is also why the term GEO carries a citation anchor its rivals lack — the terminology comparison covers that.
The keyword-stuffing result is the one worth dwelling on: it didn't just fail to help, it underperformed doing nothing. That's a genuinely useful negative finding, and it's the kind of result this scoreboard exists to preserve rather than let quietly drop out of circulation as the study ages. The same preservation problem applies across the statistics hubs generally.
Keyword stuffing didn't just fail to improve GEO citation rate in the one controlled study that's tested it — it underperformed doing nothing at all.
Share on XHow the Princeton study actually worked
This single study carries so much evidentiary weight for the field that it's worth describing its method in more detail than a single citation typically allows. The researchers constructed a benchmark of roughly 10,000 real-world queries across multiple topic domains. They then tested a set of content-level interventions: adding quotations, adding statistics, citing sources, restructuring for readability, and, as a deliberate negative control, keyword stuffing, against a generative-answer visibility metric of their own design.
Each intervention was applied to otherwise-comparable content, and the resulting visibility lift measured relative to an unmodified baseline. This is what makes the study's grade of Fact defensible on this scoreboard: a disclosed sample, a disclosed method, peer review, and a built-in negative control that could have shown no effect, or a positive one, but instead showed a clear underperformance. That's exactly the kind of result that's hard to produce by accident or motivated reasoning. Practitioner write-ups on LLM-friendly content and why original research wins citations describe the same mechanism without testing it.
The schema exception
Schema markup is graded Open question rather than settled in either direction, because the actual evidence conflicts:
Ahrefs tracked 1,885 pages that added schema markup and found AI citations "barely moved" — observational, but a real, disclosed test.
Separate live-fetch tests across five major AI systems found that none of them used information present only in JSON-LD when fetching a page directly — every system extracted visible HTML only.
Counter-argument: schema may still assist indexing or retrieval stages that happen before the live fetch — this hasn't been ruled out, only the direct-fetch mechanism has been tested against. Google's own AI optimization guide does not settle it either way, and whether crawlers even render the page is a separate open question.
This is exactly the kind of question a randomized controlled test could resolve, and it's registered as one of the pre-registered studies in the dedicated schema RCT, part of the AI Citation Index's research program.
What actually moves the needle, on current evidence
If you can only act on the tactics graded Fact or Evidence : cite your sources, include direct quotations and statistics rather than paraphrased claims, publish original data where you can, and don't rely on schema alone to do work that content structure and citation earn. The tactical version of that, written out, is how to get cited by ChatGPT; the umbrella discipline is AI SEO. Everything else on the list is a reasonable bet, not a proven lever — which is a different thing to tell a client than "this is a ranking factor."
What this scoreboard argues against, explicitly
Two specific bad habits this page is designed to push back on. First: presenting a Hypothesis-grade tactic to a client or in published content as though it carries Fact-grade certainty. That's the single most common inflation in GEO advice, and the one this scoreboard exists specifically to correct. Second: dismissing a Hypothesis-grade tactic as worthless because it lacks a controlled study. Untested isn't the same as disproven. Several of these tactics remain reasonable to implement for reasons entirely separate from AI citation. Fresh content and clear internal linking both plausibly help classic SEO and user experience, regardless of any GEO effect.
How this differs from classic SEO ranking-factor lists
Classic SEO has decades of large-scale correlation studies, and occasional controlled tests, behind its own ranking-factor discourse. That discourse is itself imperfect, but it rests on a much larger evidence base than GEO currently has. A useful contrast: Google's own ranking system incorporates hundreds of signals refined over two decades of continuous testing at massive scale, observed indirectly through correlation studies and semi-official guidance — including on how long ranking itself actually takes, a question with no AI-citation equivalent except the crawl-to-citation latency study.
GEO, by comparison, has one controlled academic study and a handful of vendor observational reports. It's a field still in its first few years, where the honest answer to "does X affect AI citation" is far more often "we don't know yet" than current content in the space tends to admit. The wider picture is the state of AI search.
How this updates
Re-graded quarterly, or immediately when a new controlled study is published. A tactic's grade moving (up or down) will be logged with the date and the study that caused the change — the same discipline as every other reference asset on this site.
Last verified: September 2026
- All 14 rows reviewed against the evidence located as of this date. No grade changed.
- Added a per-row regrade date, so a stale grade shows as a stale date rather than hiding behind a page-level "last updated".
How a tactic gets graded, step by step
The grades are the whole product here, so the procedure behind them should be visible. It is deliberately mechanical.
Step one. Find the strongest public evidence for the tactic. Not the most cited article. The strongest study.
Step two. Check whether the method is disclosed. Sample, dates, queries, engines. If those are missing, the study cannot support a grade above Hypothesis, regardless of how large the claimed effect is.
Step three. Check whether the study manipulated the variable or merely observed it. Observation caps the grade at Evidence. Only manipulation can support a causal reading.
Step four. Check independence. A study by a party that sells the tactic is not disqualified, but it is discounted, and the conflict is recorded.
Step five. Check that the outcome measured is citation, not a proxy. Crawl volume, impressions and rankings are all different outcomes. Substituting one for another is the most common inflation we encounter.
Step six. Record the grade with the date and the specific source. A grade without a source is an opinion wearing a badge.
Applying this consistently is why so many rows land on Hypothesis. That distribution is the finding, not a shortcoming of the search.
- 1 Find the strongest evidence Not the most-cited article. The strongest study.
- 2 Check disclosure Sample, dates, queries, engines. Missing any of them caps the grade at Hypothesis.
- 3 Manipulated or observed? Observation caps the grade at Evidence. Only manipulation supports a causal reading.
- 4 Check independence A study by a party selling the tactic is not disqualified, but it is discounted and the conflict recorded.
- 5 Check the outcome Citation, not a proxy. Crawl volume, impressions and rankings are different events.
- 6 Record grade, date, source A grade without a source is an opinion wearing a badge.
Evidence is not the only axis: cost and reversibility
A grade tells you how much to trust a claim. It does not tell you whether to act. Two more axes decide that, and they are often more decisive.
| Axis | Question to ask | Why it matters |
|---|---|---|
| Cost | How much work is this, in hours you actually have? | A weakly evidenced tactic that costs an hour is a cheap bet. The same tactic at three months is not. |
| Reversibility | Can you undo it if it does nothing? | Reversible changes let you act on weak evidence responsibly. Irreversible ones do not. |
| Independent value | Does it help anything besides AI citation? | Several tactics here are defensible for readers or classic search alone. That makes the GEO question moot. |
| Downside risk | What happens if it backfires? | One tactic on the board underperforms doing nothing. Others could plausibly harm readability. |
| Measurability | Will you know if it worked? | An unmeasurable tactic never leaves its current grade, no matter how long you run it. |
Reading the scoreboard on evidence alone produces a mistake we see often. Teams skip a cheap, harmless, weakly evidenced change while spending weeks on a strongly evidenced one that does not fit their content. Both axes belong in the decision.
Confounds that inflate almost every tactic claim
Most published GEO evidence is observational. That is not automatically bad, but it carries a standard set of confounds. Knowing them changes how you read every row above.
Selection. Pages that adopt a tactic differ from pages that do not, in budget, in editorial care, in authority. The tactic may be a marker of a good team rather than a cause of anything.
Simultaneous changes. Almost nobody makes one change in isolation. A relaunch that adds structured data usually also rewrites the content. Attribution to one component is guesswork.
Time drift. Engines change during the study window. A before-and-after comparison can measure the engine's evolution rather than the tactic's effect.
Survivorship. Case studies of tactics that worked are far easier to publish than case studies of tactics that did nothing. The visible record is filtered before you read it.
Outcome substitution. Reporting a rise in crawl requests as evidence for citation is common and wrong. They are different events and can move in opposite directions — crawler statistics measures the first, citation the second. Clicks are a third thing again: see AI search CTR and Ahrefs on AI Overviews reducing clicks.
Hypothesis We think selection and simultaneous changes account for a large share of the gap between claimed and real effect sizes in this field. That is a reasoned suspicion, not a measured one.
What does not transfer between engines
A tactic that works on one system may do nothing on another. The scoreboard is deliberately engine-agnostic, which is a simplification worth naming.
| Does not transfer | Why not |
|---|---|
| Effect sizes | Different retrieval stacks weigh candidate sources differently. The magnitude is stack-specific. |
| The role of conventional rankings | Some surfaces draw candidates from a search index. Others fetch live. Ranking matters enormously in one case and little in the other — see the top-10 overlap work and its sharp decline. |
| Structured data handling | Whether a system reads markup at fetch time varies. A tactic that depends on it inherits that variance. |
| Freshness sensitivity | An engine that caches aggressively will not reward an update the way a live-fetching one might. |
| Query coverage | Tactics tested on informational queries may not hold for commercial or navigational ones — YMYL and commercial queries each behave differently, and query fan-out (patent, corpus study) splits them further. |
Where a per-engine breakdown exists, we would rather publish it than a blended grade. It does not exist for most of these rows, and we say so rather than smoothing over it. The per-engine citation-statistics pages — ChatGPT, Perplexity, Claude, Gemini, AI Overviews, AI Mode and Bing Copilot — are where any such breakdown would land, along with the source delta study. Vendor studies such as Semrush's AI Overviews work and Ziptie's ChatGPT teardown each cover one surface only.
How to test one tactic on your own site
Site-level testing is imperfect, but it beats accepting a grade on faith. Here is a procedure that avoids the worst mistakes.
Step one. Pick one tactic. One. Testing three at once produces an uninterpretable result.
Step two. Split your pages into two matched groups. Match on topic, length and age, not at random from a small pool. Small random splits go lopsided easily.
Step three. Write down your prediction and your success criterion before you start. This single step eliminates most self-deception.
Step four. Apply the change to one group only. Leave the other alone completely, including the parts you are tempted to improve.
Step five. Sample citations for a fixed query set on a fixed schedule. Same queries, same days, same engines. Consistency matters more than volume.
Step six. Run it long enough for recrawl to happen. Too short a window measures crawl latency, not the tactic.
Step seven. Report the result whichever way it lands, including to yourself. A test you only believe when it agrees with you is not a test.
Expect a noisy answer. Site-level samples are small, and citation is a volatile outcome. A null result here is common and genuinely informative — the two first-party log studies published here, zero to cited and the llms.txt study, both work this way.
A worked before-and-after page
To make the top-graded tactics concrete, here is an invented illustration. The numbers are placeholders, not measurements.
Before. A paragraph reads: "Adoption of this technology has grown rapidly in recent years, and most experts agree it will continue." No source, no figure, no attributable claim.
After. The same paragraph names a specific finding, attributes it to a named study with a date, and links it. It states the sample size. It quotes one sentence from the source directly.
What changed structurally is that the passage is now quotable with attribution. A system composing an answer can lift a claim from it and say where the claim came from.
Notice what did not change. The paragraph is not longer by much. No markup was added. No keyword was repeated. The intervention is editorial, and it is the kind of change the strongest evidence on this board supports.
Evidence We present this as an illustration of the mechanism, not as a demonstration of an effect. We did not run this rewrite as a controlled test, and we are not implying we did.
Common misreadings of this scoreboard
"Hypothesis means it does not work." It means nobody has tested it in public. That is a statement about the literature, not about reality.
"The grades rank the tactics by value." They do not. They rank the evidence. The cost and reversibility section exists because value needs other inputs.
"Effect sizes are additive." Applying three tactics does not sum their reported lifts. The studies did not test combinations, and overlapping mechanisms almost certainly do not stack cleanly.
"A grade applies everywhere." It applies to the systems and content types the underlying study covered. Extending it further is our extrapolation, and readers should discount accordingly.
"Untested means unsafe to recommend." Only if you present it as proven. Recommending a cheap, reversible change while stating the uncertainty is entirely defensible. Page speed is the clearest case: Core Web Vitals and their published thresholds are well-evidenced for users and classic search, and entirely untested for AI citation. Both statements are true at once.
Null results we would publish
Every planned test behind this board can fail, and several probably will. Here is what we would do with each.
If a registered tactic test finds nothing, the row moves to Open with the date and the study. No softening, no "inconclusive but promising".
If a replication contradicts a top-graded row, that row is downgraded even though it weakens the page's most quotable content.
If our own test design proves unworkable, we publish the design failure. Knowing a measurement cannot be made cleanly is worth publishing on its own.
If no tactic separates from noise at all, we say that. A board where nothing is demonstrable is a real possible state of the field. The llms.txt row is the template: the proposal, an adoption study, a request-tracking study, a count showing 97% of files get zero AI requests, a 300,000-domain analysis finding no clear effect, and Google calling it speculative. That is what a well-evidenced negative looks like.
Reporting a Hypothesis grade to a client
The hardest conversation this page creates is telling someone that most of the advice they have been given is untested. Here is a way through it.
Separate the recommendation from the confidence. "I recommend doing this, and the evidence behind it is weak" is a coherent sentence. Most advice in this field collapses those two into one.
Lead with the cheap, reversible items. They let you make progress without staking credibility on an unproven claim.
Be explicit that you are not promising a number. If you quote an effect size from a study, name the study and its date, and say it may not transfer.
Offer a measurement plan rather than a guarantee. "Here is how we will know" is more valuable to a serious client than a forecast, and it survives contact with reality.
Finally, say what you will do if it does not work. Naming the exit in advance is the strongest available signal that the recommendation is honest.
Who this scoreboard is for
Useful to you if you are deciding where to spend limited content or engineering time, or if you have to defend a recommendation to someone sceptical.
Less useful if you are looking for a checklist to execute without judgement. The whole design of the page resists that, because the grades vary too much for a flat list.
Not useful at all if your site has a more basic problem. If AI crawlers cannot reach or read your pages, no tactic on this board applies yet. Fix access first: the bot registry names every crawler and the technical GEO audit works through the checks in order.
What would change this page
- Any new controlled study on a listed tactic, in either direction. Grades move on evidence, not on how long a row has held.
- A replication of the strongest existing result on current systems, which would either extend or expire its effect sizes.
- A vendor publishing raw data behind an observational claim, which could lift several rows from Hypothesis to Evidence.
- Evidence that tactic effects differ sharply by engine, which would force the board to split into per-engine columns.
- A demonstration that combined tactics interact, which would undercut the row-by-row structure entirely.
Open questions
Open question Do these tactics interact? Nothing in the public record tests combinations, and the row-by-row format quietly assumes independence.
Open question How quickly does an effect decay as a tactic becomes common? If everyone cites sources well, the differentiator disappears. No published work addresses this.
Open question Are effects different for small sites than large ones? Most available data comes from established domains.
Open question Does any tactic help on one engine while hurting on another? We have no data either way, and it is the scenario that would most complicate every recommendation here.
Limitations
- The Princeton study is aging relative to how fast these platforms change — see the caveat in §2's FAQ and §6's detail.
- "No study located" is not the same as "no effect exists." Several Hypothesis-grade tactics may well work; they simply haven't been tested in anything we could find.
- Grades reflect public evidence only. Vendors likely hold internal data that could shift several of these grades if published.
- This list is not exhaustive. It covers the fourteen most commonly recommended tactics encountered during this site's research, not every GEO tactic ever proposed. Newer agent-facing surfaces — agentic browsers, MCP servers and WebMCP — have no tactics on the board yet because nothing about them has been tested.
- Where sources concentrate is a separate question from which tactics work — see the most-cited domains. Everything on this board feeds the AI Citation Index, listed on the studies index.
Next: pick the single highest-graded tactic that applies to you and test it on your own site — one change, one frozen question set, measured before and after. Establish the "before" first with the AI Overview checker. If a tactic on this board comes back null for you, that result belongs in the same place ours do: the null-results registry.
Namdev, R. (2026). The GEO Tactic Evidence Scoreboard (v3). Retrieved from https://ritiknamdev.com/blog/geo-tactic-evidence-scoreboard Published under CC BY 4.0 — reuse freely with attribution.
For the study these grades draw most heavily on, see the Princeton GEO paper referenced below. For a worked example of a tactic tested to a null result, see the planned schema RCT registered in the Citation Index.