Synthesis · September 2026 snapshot

The state of AI search: September 2026

A snapshot of where AI-search research actually stands as this site's research program launches — five real findings, five genuine unanswered questions, and an honest inventory of what's been published.

Ritik Namdev Ritik Namdev ·Published September 2026 ·49 research pages published ·13 min read
The short version

As of September 2026, AI-search research has an evidence problem, not primarily a content problem: a handful of genuinely rigorous studies (Ahrefs, the Princeton GEO paper) carry most of the field's reliable findings, while most circulating statistics are vendor-reported from private corpora or, in several documented cases, simply don't reconcile with each other. This snapshot inventories what's solid, what's genuinely open, and what this site has built toward closing that gap so far.

Why a snapshot now

This site's research program, the AI Citation Index and everything built around it, launched with a specific thesis: the field is saturated with advice and starved of evidence. Fifty research pages later, that thesis holds up. It's worth stating plainly, with receipts, rather than as an abstract claim made once at the start and never revisited.

The landscape, in five findings

Evidence

Google ranking no longer strongly predicts AI citation. Overlap ranges from ~12% (ChatGPT/Gemini/Copilot) to ~38% (AI Overviews, down from ~76% a year earlier) — see most-cited domains and AI Overview statistics.

Evidence

Engines barely agree with each other. Reported ChatGPT/Perplexity domain overlap sits around 11% — see cross-platform concordance.

Evidence

The schema debate has real evidence on both sides, and it's unresolved. Observational data shows no clear lift; live-fetch tests show schema often ignored entirely — see the schema RCT design.

Fact

Three separate studies collapsed into one "58%" figure that gets misapplied across contexts — the clearest documented case of statistic conflation in the field, per the provenance audit.

Open question

Not one randomized, causal test of a GEO tactic existed in public before this year's pre-registrations. Every published finding in the field prior to this site's Tier 3 program is correlational.

Composition of this site's 50-page research launch
  • Statistics & platform pages 20
  • Pre-registered studies 15
  • Reference & framework 9
  • Definitions & glossary 6
First-party count, September 2026 — statistics/platform pages, pre-registered studies, reference/framework assets, and definitions.

What these five findings mean together

Read individually, each finding above is a discrete fact about one narrow slice of AI search. Read together, they describe something more structural. The traditional shortcut, rank well on Google, get cited by AI, no longer reliably works. A strategy tuned for one engine doesn't transfer to the next.

Even a widely recommended technical tactic, schema, has genuinely mixed evidence, not a settled answer. Sloppy statistic handling is common enough to produce a documented three-studies-collapsed-into-one case. And almost nothing has been tested with a method that could actually support a causal claim. Put plainly: the practical implication isn't "GEO doesn't work." It's "almost nobody has actually tested what works." That's precisely the gap this site's research program exists to close.

Fifty pages into building an independent AI-search research site, the finding holds: the field has an enormous amount of published conclusion and very little published evidence.

Share on X

The five biggest unanswered questions

  1. What does Claude actually cite? The least-measured major surface — see Claude citation statistics.
  2. How many sub-queries does AI Mode really generate per question? Folklore says 8–16; nobody has published the data — see the Fan-Out Corpus design.
  3. Does anything a site controls causally change its citation rate? Nearly everything published is correlational — see the Tier 3 RCT program (schema, freshness, author bio).
  4. How long does a citation last? Every study is a snapshot; none track duration — see citation half-life.
  5. Do AI crawlers render JavaScript? A binary, foundational, and still-untested technical question — see the rendering experiment.

What surprised us most while building this

Two things stood out more than expected going into this project. First, how often a specific, precise-sounding statistic, "47.9% of ChatGPT citations are Wikipedia," turned out, on tracing, to originate from exactly one vendor-reported source, repeated verbatim across dozens of secondary articles. The appearance of broad consensus masked a single underlying data point.

Second, how genuinely thin the causal-research layer is. Not a niche gap in one corner of the field, but essentially the entire field, across every major platform and tactic, prior to the pre-registrations this site has begun publishing this year. Neither is a criticism of any individual source. Most are transparent about being vendor-reported. But the aggregate picture, once actually traced hop by hop, is thinner than the sheer volume of published content about AI search would suggest.

What this site has published so far

Forty-nine research pages as of this snapshot. The Citation Index's pre-registration and its core metrics. Thirteen curated and cross-platform statistics pages. Ten pre-registered original studies. Six Claude-Code-for-SEO practitioner pages. And a reference layer: the AI Bot Registry, the tactic scoreboard, the Null-Results Registry, the glossary, and two definitional pillars.

Methodology lessons from the first fifty pages

A few practical lessons worth stating plainly, since they shaped how later pages on this site were built. Tracing a statistic to its actual origin routinely takes longer than writing the page around it. It's worth doing anyway, since it's the entire basis for this site's credibility claim.

A pre-registered hypothesis is only meaningful if the prediction is genuinely stated before data collection, not written to match a result decided in advance. That's why several studies here register a directional prediction that could turn out wrong, rather than a safe, hedge-everything framing.

And a null-result commitment only means something if an actual null result eventually gets published, not merely stated as an intention. That's exactly why the Null-Results Registry exists as a standing, checkable mechanism, not a one-time promise.

What comes next

Data collection for Index v1 begins Q1 2027. Every pre-registered study above moves from design to result on that same timeline. The next State of AI Search edition, in Q4 2027, will report actual findings against every prediction registered here — including the ones we predicted would fail.

How this snapshot was built

A first-party inventory and synthesis of this site's own published research as of the date shown, cross- referenced against the provenance grades established throughout. No new data collection underlies this specific page — it's a synthesis of work already published and linked above.

How to cite this
Namdev, R. (2026). The state of AI search: September 2026 (v1). Retrieved from https://ritiknamdev.com/blog/state-of-ai-search-september-2026

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Every finding and gap below links to its own dedicated page. Start with the AI Citation Index if you're new to this research program.

FAQ

Frequently asked questions

Is this the annual "State of AI Search" report described in the site's strategy?
This is an early snapshot, not the full annual report. That's scheduled for Q4 2027, once a full year of Citation Index data exists. This page marks where the research program stood as of September 2026, at the site's own launch of this research push.
What's the single most important finding across everything published so far?
That the field has an enormous amount of published conclusion and very little published evidence. Nearly every genuinely rigorous, disclosed-method study traces to a small handful of sources, Ahrefs, the Princeton paper. Most circulating statistics are vendor-reported from non-public corpora, or, in a few cases, don't reconcile with each other at all.
How will this snapshot be updated?
It gets superseded by later, fuller editions as the Citation Index accumulates real data. This specific page stays as a dated historical marker, rather than being silently rewritten.
Why publish a snapshot at all before any first-party data exists?
Because the honest starting state matters as a baseline. A later edition claiming progress is only checkable against a plainly stated account of where things stood before that progress happened. Publishing this now, warts and gaps included, is what makes the eventual Q4 2027 comparison meaningful, rather than retroactively flattering.
Does this snapshot represent a consensus view of the field, or just this site's perspective?
This site's own perspective and inventory specifically. The findings and gaps are grounded in what we could verify through direct source-tracing, not a survey of practitioner opinion or an attempt to represent every viewpoint in the field.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.