As of September 2026, AI-search research has an evidence problem, not primarily a content problem: a handful of genuinely rigorous studies (Ahrefs, the Princeton GEO paper) carry most of the field's reliable findings, while most circulating statistics are vendor-reported from private corpora or, in several documented cases, simply don't reconcile with each other. This snapshot inventories what's solid, what's genuinely open, and what this site has built toward closing that gap so far.
Why a snapshot now
This site's research program, the AI Citation Index and everything built around it, launched with a specific thesis: the field is saturated with advice and starved of evidence. Fifty research pages later, that thesis holds up. It's worth stating plainly, with receipts, rather than as an abstract claim made once at the start and never revisited. Whether the discipline gets called GEO, AI SEO or one of the competing acronyms makes no difference to that diagnosis.
The landscape, in five findings
Google ranking no longer strongly predicts AI citation. Overlap ranges from ~12% (ChatGPT/Gemini/Copilot) to ~38% (AI Overviews, down from ~76% a year earlier) — see most-cited domains, AI Overview statistics and Ahrefs' underlying measurement.
Engines barely agree with each other. Reported ChatGPT/Perplexity domain overlap sits around 11% — see cross-platform concordance. The same divergence is described qualitatively by Discovered Labs, Leapd and Profound.
The schema debate has real evidence on both sides, and it's unresolved. Observational data — Ahrefs' correlational look included — shows no clear lift; live-fetch tests show schema often ignored entirely — see the schema RCT design. The llms.txt story is the same shape: 300,000 domains, no clear effect, and Google calling it speculative.
Three separate studies collapsed into one "58%" figure that gets misapplied across contexts — the clearest documented case of statistic conflation in the field, per the provenance audit.
Not one randomized, causal test of a GEO tactic existed in public before this year's pre-registrations. Every published finding in the field prior to this site's Tier 3 program is correlational.
of AI citations on ChatGPT, Gemini and Copilot (grouped) also rank in Google's top 10 for the same query.
domain overlap reported between ChatGPT's and Perplexity's citation sets — two engines answering the same questions from largely different sources.
randomized, causal tests of a GEO tactic existed in public before this year's pre-registrations.
- Statistics & platform pages 20
- Pre-registered studies 15
- Reference & framework 9
- Definitions & glossary 6
What these five findings mean together
Read individually, each finding above is a discrete fact about one narrow slice of AI search. Read together, they describe something more structural. The traditional shortcut, rank well on Google, get cited by AI, no longer reliably works — even though Seer found one engine matching Bing's results at 87%, which is exactly the kind of single-engine result that does not generalise. A strategy tuned for one engine doesn't transfer to the next.
Even a widely recommended technical tactic, schema, has genuinely mixed evidence, not a settled answer. Sloppy statistic handling is common enough to produce a documented three-studies-collapsed-into-one case. And almost nothing has been tested with a method that could actually support a causal claim. Put plainly: the practical implication isn't "GEO doesn't work." It's "almost nobody has actually tested what works." That's precisely the gap this site's research program exists to close.
Fifty pages into building an independent AI-search research site, the finding holds: the field has an enormous amount of published conclusion and very little published evidence.
Share on XThe five biggest unanswered questions
- What does Claude actually cite? The least-measured major surface — see Claude citation statistics.
- How many sub-queries does AI Mode really generate per question? Folklore says 8–16; nobody has published the data, though Google has described the mechanism — see the Fan-Out Corpus design, and the source delta between AI Mode and AI Overviews.
- Does anything a site controls causally change its citation rate? Nearly everything published is correlational — see the Tier 3 RCT program (schema, freshness, author bio).
- How long does a citation last? Every study is a snapshot; none track duration — see citation half-life.
- Do AI crawlers render JavaScript? A binary, foundational, and still-untested technical question — see the rendering experiment, and the bot registry for which crawlers would be affected.
What surprised us most while building this
Two things stood out more than expected going into this project. First, how often a specific, precise-sounding statistic, "47.9% of ChatGPT citations are Wikipedia," turned out, on tracing, to originate from exactly one vendor-reported source, repeated verbatim across dozens of secondary articles. The appearance of broad consensus masked a single underlying data point — the same pattern behind the Wikipedia dependency figure and Perplexity's Reddit share, and the reason the statistics pages here carry a verification-status column.
Second, how genuinely thin the causal-research layer is. Not a niche gap in one corner of the field, but essentially the entire field, across every major platform and tactic, prior to the pre-registrations this site has begun publishing this year. Neither is a criticism of any individual source. Most are transparent about being vendor-reported. But the aggregate picture, once actually traced hop by hop, is thinner than the sheer volume of published content about AI search would suggest.
What this site has published so far
Forty-nine research pages as of this snapshot. The Citation Index's pre-registration and its core metrics. Thirteen curated and cross-platform statistics pages. Ten pre-registered original studies. Six Claude-Code-for-SEO practitioner pages, with their failure modes and open tooling. And a reference layer: the bot registry, the tactic scoreboard, the Null-Results Registry, the glossary, and two definitional pillars. Everything is listed under studies, with tools alongside.
Methodology lessons from the first fifty pages
A few practical lessons worth stating plainly, since they shaped how later pages on this site were built. Tracing a statistic to its actual origin routinely takes longer than writing the page around it. It's worth doing anyway, since it's the entire basis for this site's credibility claim.
A pre-registered hypothesis is only meaningful if the prediction is genuinely stated before data collection, not written to match a result decided in advance. That's why several studies here register a directional prediction that could turn out wrong, rather than a safe, hedge-everything framing.
And a null-result commitment only means something if an actual null result eventually gets published, not merely stated as an intention. That's exactly why the Null-Results Registry exists as a standing, checkable mechanism, not a one-time promise.
What comes next
Data collection for Index v1 begins Q1 2027. Every pre-registered study above moves from design to result on that same timeline. The next State of AI Search edition, in Q4 2027, will report actual findings against every prediction registered here — including the ones we predicted would fail.
- Sep 2026This snapshot
49 research pages published; no first-party citation data yet.
Published deliberately before any results exist, so a later claim of progress is checkable against a stated baseline.
- Q1 2027Index v1 collection
Data collection begins for the Citation Index and every registered study.
Design to result on one shared query set and one collection window.
- Q4 2027Full annual report
Findings reported against every prediction registered here.
Including the predictions we expect to fail. That is the point of registering them.
A worked example: reading one claim end to end
Take a claim of the shape you meet every week. "AI answers cite forums more than brand sites." How should you read it? The steps below are the same ones this site applies internally.
First, ask what was counted. A linked source, or any mention of a name? The two definitions produce very different tables from the same set of answers.
Second, ask which questions were asked. Forums do well on troubleshooting and opinion. If the query set leans that way, the finding is about the query set, not about forums in general.
Third, ask when. A surface that favoured forums one quarter may not the next. Without a date the claim cannot be checked against anything.
Fourth, ask who measured it and what they sell — the same question worth asking of a 23-fold conversion claim or a brand-mentions visibility gap. This is not an accusation. It is context that tells you which direction an unconscious thumb might press.
Fifth, ask what the claim would predict for you. If it is true, you should be able to see something in your own data. If it predicts nothing observable, it is not actionable regardless of whether it is true.
Run those five questions and most claims resolve into one of three states. Checkable and probably sound. Plausible but unverifiable. Or meaningless as stated. Sorting into those three buckets is most of the work.
A field guide to the evidence types
Most confusion in this field comes from treating different kinds of claim as equivalent. They are not. Here is the ladder, from weakest to strongest.
| Type | What it can support | What it cannot |
|---|---|---|
| Anecdote | A hypothesis worth testing | Any general claim |
| Vendor aggregate, private corpus | A description of that vendor's sample | A claim about your site or the web |
| Observational study, disclosed method | An association, with stated limits | Cause |
| Live-fetch test | What a system did on a given day | What it will do next month |
| Pre-registered randomised test | A causal estimate within its scope | Transfer to a different engine or vertical — local, commerce and YMYL queries all behave differently |
Notice the right-hand column. Every row has one. Even the strongest design has a boundary, and naming it is not modesty. It is what makes the finding usable.
Why the evidence layer is thin
This is not a story about laziness. The incentives explain most of it.
The data is the product. A visibility vendor's corpus is what customers pay for. Publishing raw files would give it away. So the method stays private, and the claim cannot be checked.
Null results do not market. A finding that a tactic did nothing sells no software and wins no links — which is why the Null-Results Registry exists as a standing commitment rather than a good intention. The llms.txt literature is the exception that proves it: adoption tracking and the finding that most files are never requested got far less circulation than the original proposal did. The publishing filter therefore selects for positive results.
Causal tests are expensive. A randomised design needs many comparable pages, a control group, and the willingness to leave the control group alone for months. Most sites will not do that.
The systems move. A result can be obsolete before it is written up. That discourages careful work and rewards fast commentary.
Nobody is required to disclose anything. There is no registry, no peer review, and no convention that a number should carry its method — the gap the measurement standard and the dataset strategy are both aimed at. So most numbers do not.
Hypothesis Our reading is that these incentives, not any shortage of skill, explain the shape of the field. We cannot prove that, and a different explanation may fit the same facts.
How to trace a statistic yourself
This is the single most useful habit in the field, and it takes about fifteen minutes per number.
- 01 Copy the exact figure Include the decimal. A precise number is easier to follow through a citation chain than a rounded one.
- 02 Find the earliest carrier Search the figure in quotes, sort by date, ignore anything that merely links onward.
- 03 Follow every hop A cites B cites C. Keep going until you reach a page claiming to have measured something.
- 04 Interrogate the origin Sample? Dates? Which engine? Which country and device mix? Is the data available?
- 05 Record the answer Including “none given”. A missing method is a result. Write it beside the figure.
- 06 Check for meaning drift This is where conflation happens — a number measured on one thing restated about another.
1. Copy the exact figure. Include the decimal. A precise number is easier to follow through the citation chain than a rounded one.
2. Find the earliest page that carries it. Search the figure with quotation marks. Sort by date if you can. Ignore anything that merely links onward.
3. Follow every hop. Article A cites B, which cites C. Keep going until you reach a page that claims to have measured something, not one that quotes.
4. At the origin, ask five questions. What was the sample? Over what dates? Which engine? Which country and device mix? Is the underlying data available?
5. Record the answer, including "none given". A missing method is a result. Write it down next to the figure so you never have to repeat the trace.
6. Check whether the figure changed meaning along the way. This is where conflation happens. A number measured on one thing gets restated about another, and nobody notices.
What does not transfer between engines
Low cross-engine overlap has a practical consequence that is often skipped. A finding about one engine is a finding about that engine only.
Retrieval differs. One system may lean on a live search index. Another may lean on its own crawl, or on material already inside the model. These produce different source sets from the same question.
Answer format differs. A system that shows numbered citations creates different click behaviour from one that mentions a brand in prose without a link.
Freshness policy differs — a registered experiment here, and separately how long a citation survives at all, measured in the half-life study. Some surfaces clearly prefer recent material. Others repeat the same source for months. A freshness result on one is not a freshness result on all.
User population differs — Similarweb's category usage data, Statista's ChatGPT user series and StatCounter's search share each describe a different population, as the market share page explains. The questions people bring to each product are not the same, so the query mix behind any aggregate is different too.
The rule this site follows is to report per engine, and to treat any blended figure as a description of the blend rather than of AI search in general.
Confounds in reading the landscape
Even careful readers can draw the wrong conclusion from a correct set of facts. A few confounds recur.
Survivorship in what gets published. You see the studies that found something. The ones that found nothing were usually never written up.
Sampling by convenience. Query sets are often drawn from whatever the researcher had access to. That is not a random sample of what people ask.
Time compression. Findings from different quarters get quoted side by side as though they describe the same system. These products change substantially between quarters.
Definition drift. "Citation" means a linked source in one study and a brand mention in another. Compared directly, the two produce nonsense.
Popularity bias. Domains that appear often in AI answers are usually large and old. Attributing their presence to a tactic ignores everything else that makes them prominent.
Common misreadings of this snapshot
"The site says GEO does not work." It does not say that. It says the public evidence for most specific tactics is weak or absent. Those are different claims.
"Vendors are lying." Most are transparent that their figures are vendor-reported. The problem is downstream, where the qualifier gets dropped and the number becomes a fact.
"Low engine overlap means AI search is random." It means the systems differ. Each may be internally consistent while disagreeing with the others.
"Wait for the evidence before doing anything." Also wrong. Fundamentals such as accessible content, clear answers and a crawlable site are defensible on their own terms — as is the reading of original research as a citation asset that the citation playbook builds on.
"This snapshot is the field's consensus." It is one site's traced inventory on one date. It is meant to be argued with, and it names the date so it can be.
A field precedent: early SEO went through this
Search optimisation in its first decade looked much like this. Confident claims. Private data. Ranking factor lists — the modern descendant being correlational AI citation factor rankings — assembled from correlation and repeated until they felt like physics. Even the more careful statistics roundups inherit that lineage, and the encyclopedic account of the field rests on a single academic paper (PDF).
What changed it was slow and unglamorous. Public correlation studies with stated methods. Controlled tests that people could replicate. Platform documentation that could be checked against behaviour. And a culture that eventually treated an untraceable number as embarrassing rather than persuasive.
The precedent is encouraging in one way and cautionary in another. Encouraging, because the shift did happen. Cautionary, because it took years, and plenty of folklore from the early period survived long after it was disproved.
Hypothesis We expect AI search to follow a similar arc, faster, because the practitioner community already knows what a disclosed method looks like. That expectation is a guess, not a forecast grounded in data.
Who this snapshot is for
It is for someone who has to make a decision and wants to know how much weight the available evidence can bear. That includes in-house teams, consultants and anyone writing about the field.
It is useful if you are being asked to justify a budget — alongside referral traffic, click-through, zero-click and conversion figures, all of which carry the same caveats as everything else here. The honest answer, that the causal evidence is thin and here is what we will measure instead, is a stronger position than a borrowed statistic.
It is less useful if you want a checklist. This page deliberately does not provide one, because a checklist would imply a confidence the evidence does not support.
It does not apply if your traffic and revenue do not touch these surfaces at all. Some businesses will see no meaningful volume from AI answers for a long time, and that is a legitimate finding for them.
The cost of acting on weak evidence
Weak evidence is not free. It has predictable costs, and they compound.
Wasted implementation. Rolling a tactic across thousands of pages because of an unverified figure consumes engineering time that had an alternative use.
Damaged credibility. When the number turns out to be untraceable, everything else you said gets re-examined. That cost lands on the person who repeated it, not the vendor.
Lost learning. If you change five things at once, you learn nothing about any of them. The next decision is as blind as the last.
Blocking the wrong thing. Access decisions have the same evidence problem — the blocking census, the training-versus-retrieval distinction and BuzzStream's publisher study all describe policies set on assumption rather than measurement.
Wrong attribution. Acting on a bad model leads you to credit or blame the wrong cause, which then shapes the following quarter's plan.
The cheap alternative is to state a prediction before acting, and to write down what would count as the tactic failing. That single habit converts spend into evidence.
Reporting this to a stakeholder
Lead with what is settled. Answers are appearing above results. The sources those answers name are not simply the top-ranked pages. The systems disagree with each other.
Then be explicit about the gap. There is no public, causal evidence that a specific on-page change moves citation rate. Anyone claiming otherwise should be asked for the method.
Then propose the measurable thing. Track your own citations for a defined query set — the method the zero-to-cited study used. Track the AI-agent hits in your own logs, benchmarked against Cloudflare's 2025 crawler data (summarised here) and our own crawler statistics. Report movement, not projections.
Finally, set the review date. Say plainly that this picture is dated, and that you will bring back a revised one in a quarter. That framing protects you when the ground shifts, which it will.
Null results we would publish
- No effect from any registered tactic. If the randomised studies find nothing, that result gets the same prominence as a positive one.
- Higher engine overlap than reported. If our own measurement contradicts the low-overlap figures we have cited, we will say so and revise the finding above.
- No half-life effect. If citations turn out to be stable over time, the freshness framing loses its motivation and we will retire it.
- Our own traces being wrong. If a statistic we called untraceable turns out to have a solid origin we missed, the correction gets published with the same weight as the original criticism.
What would change this page
An operator publishing a real description of how sources are selected for an answer would change more of this page than anything else on this list.
A vendor releasing a raw, dated corpus with a documented sampling method would move several findings from vendor-reported to checkable overnight.
A large-sample observational study with a fully published corpus — the shape of Semrush's AI Overviews work or Ahrefs' click-loss measurement, but with the underlying data attached — would move several findings at once.
A second independent group running a randomised test would let us compare designs instead of relying on a single programme, which is itself a weakness of the current picture.
Any of these would be good news, and would be treated as such. This page is not invested in the field staying thin.
Open questions beyond the top five
- Does being cited in an AI answer produce measurable business value, or only visibility? Almost nothing public addresses this.
- How stable is an engine's source set for the same question asked twice, an hour apart?
- Do answers differ by country and language in ways that change which domains get named?
- What share of AI answers name no source at all, and is that share moving?
- Does a site being blocked to one crawler affect its appearance in a different company's answers?
- Does authorship or credential signalling change anything, as the registered author-bio RCT is designed to test, and do brand mentions outperform links as a correlate?
- Do agentic browsers, MCP servers and agent-readable endpoints become a separate visibility surface, or fold back into the ones measured here?
How this snapshot was built
A first-party inventory and synthesis of this site's own published research as of the date shown, cross- referenced against the provenance grades established throughout and the standards described in about. No new data collection underlies this specific page — it's a synthesis of work already published and linked above.
Namdev, R. (2026). The state of AI search: September 2026 (v1). Retrieved from https://ritiknamdev.com/blog/state-of-ai-search Published under CC BY 4.0 — reuse freely with attribution.
Every finding and gap below links to its own dedicated page. Start with the AI Citation Index if you're new to this research program.