Research programme · v0

Research & studies

Every research page on this site, in one place. The measurement programme is pre-registered and open: protocols before data, raw files alongside findings, and null results published like any other. This page says exactly what has been measured and what has not.

Where this stands, plainly
  • The AI Citation Index is at v0 — pre-registration. The protocol is published. No Index data has been collected yet.
  • 14 studies have their questions, methods and predictions published in advance, awaiting collection.
  • 56 research pages are live, most of them tracing existing public figures to their origin and grading them.
  • Two pages carry first-party observations from this site's own server logs. They are labelled as such.
  • Where a number does not exist, these pages say so rather than estimating one.

Where the programme stands

The short version: the protocols are published and the data collection has not started. That is an unusual thing to advertise, so it is worth explaining why this page leads with it.

Most AI-search statistics in circulation come from companies whose data is the product they sell. That is not an accusation of bad faith. It is a structural problem. A vendor cannot publish the raw file without giving away the thing customers pay for, and they cannot publish a null result without undermining the category they sell into.

This programme occupies the position they cannot. Protocols are published before data is collected. Raw files ship alongside findings. A result that shows nothing gets published anyway. That last commitment is the one that costs something, and it is the clearest signal that a method is real.

The honest cost of that approach is visible right here. There is less finished data on this site than on a vendor blog, because the finished data has to survive being checked.

The AI Citation Index

The signature measurement: a recurring, versioned, openly published record of what AI search engines actually cite. Every other research page on this site orbits it.

Index release status
  1. v0Sep 2026

    Protocol only — the pre-registration.

    Method, metrics and hypotheses committed before any collection.

Open question No citation data has been collected under this protocol yet. Any figure you see elsewhere on this site comes from third-party research and is graded for how far it could be traced. Read the full protocol in the AI Citation Index pre-registration, and the planned data releases in the dataset strategy.

How a study runs here

Four steps, in this order, without exception. The order is the whole point. Collecting first and deciding what counts afterwards is how an honest researcher accidentally produces a result that was never really there.

The study protocol
  1. 01 Pre-register Protocol, metrics and predictions published first
  2. 02 Collect Fixed query panel, disclosed conditions, repeated runs
  3. 03 Publish raw The dataset, not only the summary
  4. 04 Report either way A null result ships like any other

Step four is the one almost nobody commits to. Studies that find nothing are the ones that quietly disappear across every research field, which is why the published record looks more certain than the evidence is. Ours are listed in the null results registry before they run.

The collection code and analysis scripts are intended for open release, described in open-source SEO agent tooling. A finding you cannot reproduce is a claim, not a result.

Committed in advance

Pre-registered studies

Each of these has its question, method and predictions published before collection. None has results yet. Each page states what it would take to run, and what a null result would look like.

The questionStudyStatus
Do different AI engines cite the same sources for the same question? Cross-platform citation concordance Open question
Do two Google surfaces cite the same sources for the same query? AI Mode vs. AI Overviews: source delta Open question
How long does a citation survive before it is replaced? AI citation half-life: how long do citations last? Open question
How long is the gap between a crawl and a first citation? Crawl-to-citation latency Open question
Does updating a page change how often it is cited? Does content freshness increase AI citations? Open question
Does structured data change citation rates? Does schema markup increase AI citations? Open question
Do visible author credentials affect citation? Does an author bio affect AI citation? Open question
How many sub-queries does one question actually become? The Query Fan-Out Corpus Open question
What gets cited when community discussion is thin? The Reddit dependency: what Perplexity cites when Reddit has nothing Open question
What replaces the encyclopaedia when it has no entry? The Wikipedia dependency: what ChatGPT cites when Wikipedia has nothing Open question
Do unlinked mentions predict citation better than links? Brand mentions vs. backlinks: which predicts AI citation? Open question
Which AI crawlers execute JavaScript? Do AI crawlers render JavaScript? Open question
What does an agentic browser actually request? Agentic browsers: what do they actually fetch? Open question
How many sites actually block AI crawlers? The robots.txt AI-Blocking Census Open question
Traced & graded

Statistics hubs

These pages take figures that already circulate and trace each one back to its origin. Where the chain dead-ends in a vendor blog citing another vendor blog, the page says so. That observation is usually worth more than the number.

Study

The state of AI search: September 2026

A snapshot of where AI-search research actually stands as this site's research program launches. Five findings, five unanswered questions, and an honest inventory.

10 min read · Sep 2026
Study

AI search CTR statistics

How click-through behavior changes when an AI Overview appears. The specific population each figure describes stays attached, not stripped away.

14 min read · Sep 2026
Study

AI search and YMYL content: what is actually known

Health, finance, legal and safety queries carry higher stakes. Do AI engines cite differently for them? That is an open, under-studied question.

13 min read · Sep 2026
Study

Local AI search: what is actually known

Local intent is a large share of all search behavior. Rigorous AI-search research on it is almost entirely absent. Stated plainly, not filled with a borrowed number.

13 min read · Sep 2026
Study

AI shopping and agentic commerce statistics

AI agents are beginning to complete purchases on a shopper's behalf. What is actually measured, and how much of the coverage is still forecast rather than fact.

13 min read · Sep 2026
Study

Agentic browsers: what do they actually fetch?

ChatGPT Atlas, Perplexity Comet and Gemini in Chrome are shipping now. What they actually request from a page has barely been studied. Pre-registered.

13 min read · Sep 2026
Study

Do AI crawlers render JavaScript?

Retrieval bots may not execute client-side JavaScript. If not, any content that depends on it is invisible to them entirely. Pre-registered experiment.

14 min read · Sep 2026
Study

AI Mode vs. AI Overviews: source delta

Two distinct Google AI surfaces, routinely treated as interchangeable. Do they actually cite the same sources for the same query? Pre-registered study.

13 min read · Sep 2026
Study

The Reddit dependency: what Perplexity cites when Reddit has nothing

Perplexity's citations skew heavily to Reddit. For B2B and enterprise topics with thin discussion, what fills the gap? Pre-registered study.

14 min read · Sep 2026
Study

The Wikipedia dependency: what ChatGPT cites when Wikipedia has nothing

ChatGPT's citations skew heavily to Wikipedia. For topics it doesn't cover well, what fills the gap? Pre-registered study.

14 min read · Sep 2026
Study

Brand mentions vs. backlinks: which predicts AI citation?

Reported correlations put brand mentions at roughly 3x the predictive strength of backlinks for AI Overview visibility. A confound nobody has controlled for.

14 min read · Sep 2026
Study

Does an author bio affect AI citation?

Author credentials and E-E-A-T signals are recommended constantly. No controlled test isolates whether they change AI citation rate. A pre-registered RCT.

16 min read · Sep 2026
Study

Does content freshness increase AI citations?

'Keep content updated' is universal GEO advice. No controlled study has tested it against citation rate. A pre-registered RCT.

14 min read · Sep 2026
Study

Crawl-to-citation latency

How many days pass between an AI crawler first fetching a new page and that page appearing in a citation? Pre-registered panel study.

14 min read · Sep 2026
Study

The Query Fan-Out Corpus

Everyone repeats that Google AI Mode fans a query into '8-16 sub-queries.' Nobody has published the data behind that number. A pre-registered corpus design.

15 min read · Sep 2026
Study

Cross-platform citation concordance

How much do ChatGPT, Perplexity, Gemini, Claude and Google agree on what to cite? A reported ~11% two-engine overlap suggests: not much. Pre-registered study.

13 min read · Sep 2026
Study

AI citation half-life: how long do citations last?

Once a URL is cited by an AI engine, how long does it stay cited? Nobody has measured this. A 12-month tracking study, pre-registered.

15 min read · Sep 2026
Study

Does schema markup increase AI citations?

The most contested tactic in GEO. This page registers a randomized test to settle it. Pre-registered before a single page is enrolled.

16 min read · Sep 2026
Study

AI search conversion benchmarks

Published claims about AI traffic conversion range from 4.4x to 23x organic. Different orders of magnitude. Here is why, and the honest version of the statistic.

15 min read · Sep 2026
Study

AI referral traffic statistics

How much traffic AI platforms actually send to websites, how it splits between engines, and why standard analytics likely undercount it.

14 min read · Sep 2026
Study

Zero-click search statistics

The most commonly quoted zero-click figures, each attached to the specific population it actually measures. Collapsing them into one number is where most content goes wrong.

14 min read · Sep 2026
Study

AI search market share statistics

How traffic is distributed across AI search platforms, how that's shifted over the past year, and how AI referral engagement compares to classic organic.

12 min read · Sep 2026
Study

Gemini citation statistics

Gemini, AI Overviews, and AI Mode are three related but distinct Google surfaces, routinely conflated in coverage. What is actually known about Gemini specifically.

11 min read · Sep 2026
Study

Perplexity citation statistics

Perplexity is the most transparent of the major AI search engines about its sources. Its citation pattern looks meaningfully different from ChatGPT.

11 min read · Sep 2026
GEO

Claude citation statistics

What is actually known about what Claude cites, and an honest accounting of how much isn't known. Claude is the least-studied major AI search surface.

10 min read · Sep 2026
Study

ChatGPT citation statistics

What ChatGPT actually cites, how it retrieves information, and how weakly Google ranking predicts showing up in a ChatGPT Search answer.

12 min read · Sep 2026
Study

Most-cited domains in AI search

Which kinds of sources dominate citations across ChatGPT, Perplexity and Google AI Overviews. And how little the major engines actually agree with each other.

13 min read · Sep 2026
Study

Google AI Mode statistics and query fan-out, explained

How AI Mode actually retrieves information: the query fan-out mechanism, its documented patent basis, and what is confirmed versus estimated.

12 min read · Sep 2026
Study

Google AI Overview statistics

Trigger rates, citation-to-ranking overlap, and click-through impact for AI Overviews - with a verification status attached to every figure.

13 min read · Sep 2026
Study

The robots.txt AI-Blocking Census

How many of the web's most important sites block AI crawlers, measured directly from robots.txt, not quoted from a report with no visible sample. Pre-registration.

14 min read · Sep 2026
Study

Where AI SEO statistics actually come from

We traced twelve of the most-quoted numbers in AI search back to their origin. Five hold up. Three dead-end. The most-repeated figure in the field is three studies wearing one number.

22 min read · Sep 2026
Study

The AI Citation Index

A quarterly, open measurement of what seven AI search engines actually cite. Query set, raw data and methodology published in full. This is the pre-registration.

24 min read · Sep 2026
Study

Zero to Cited: how a new site climbs into AI search

The evidence-backed timeline and playbook for taking a brand-new site from launch to its first citations in Bing, Perplexity, ChatGPT and AI Overviews.

14 min read · Jul 2026
Study

Does llms.txt actually work? A 90-day log test

I deployed llms.txt and read the server logs for 90 days. Here is what the data, and the largest studies, actually say.

13 min read · Jun 2026
Interpretation

GEO research & guides

What the evidence supports, what it does not, and where the field is running ahead of its data. Start with what GEO actually is or the tactic evidence scoreboard.

GEO

The AI citation dataset strategy

Every dataset this site plans to publish, in one place: what it contains, when it ships, how it is licensed. Raw files, not just summary reports.

10 min read · Sep 2026
GEO

What is AI SEO?

The umbrella term for optimizing a site's visibility across the whole AI-search landscape: generative answers, AI crawlers, and the technical infrastructure behind both.

12 min read · Sep 2026
GEO

What is GEO?

Optimizing content so AI systems like ChatGPT, Perplexity and Gemini cite it when generating an answer. That is Generative Engine Optimization.

19 min read · Sep 2026
GEO

AI search glossary

Definitions for the terms used throughout this site's research. Kept tight, each linked to the fuller page where the concept is explored.

18 min read · Sep 2026
GEO

AI SEO statistics

Every AI-search statistic on this site, in one table, each with a verification grade and a link to full context.

17 min read · Sep 2026
GEO

AEO statistics

Answer Engine Optimization statistics. Genuinely fewer exist under this specific label than under GEO, and this page says so rather than padding itself.

5 min read · Sep 2026
GEO

GEO statistics

The most commonly cited Generative Engine Optimization statistics, each with a verification status attached rather than uniform, unearned confidence.

11 min read · Sep 2026
GEO

The Null-Results Registry

Every pre-registered prediction on this site that is, or is predicted to be, a null result, tracked in one place. Almost nobody publishes negative findings.

10 min read · Sep 2026
GEO

WebMCP and the agent-readable web

Google and Microsoft shipped an early preview of a structured agent-interaction protocol for websites in Feb 2026. What it is, and why site owners should care now.

12 min read · Sep 2026
GEO

MCP servers as a visibility channel

The Model Context Protocol lets AI agents query a business data directly. Does publishing one affect AI-search visibility? Untested, but genuinely early.

11 min read · Sep 2026
GEO

The AI Visibility Measurement Standard

'AI visibility' currently means whatever a vendor's dashboard measures. A proposed open specification so numbers from different sources become comparable.

13 min read · Sep 2026
GEO

GEO vs AEO vs LLMO: which term is winning

Three acronyms describe overlapping ideas about optimizing for AI-generated answers. The industry hasn't settled on one. What each means, where they overlap, and which is gaining ground.

18 min read · Sep 2026
GEO

The GEO Tactic Evidence Scoreboard

Every tactic the industry recommends for getting cited by AI search, graded by the actual evidence behind it. Not by how often it is repeated.

19 min read · Sep 2026
GEO

Bing Copilot SEO: the easiest AI engine to crack

ChatGPT search runs on Bing's index. Getting into Bing is the fastest path to your first AI citation. Here's the full playbook.

12 min read · Jun 2026
GEO

How to get cited by ChatGPT: 9 tactics, with data

The Princeton-backed levers - statistics, quotes, structure, freshness - that actually move AI citations.

13 min read · Jul 2026
Infrastructure

Technical research

Crawlers, access controls, rendering and the plumbing that decides whether your pages are reachable at all. Begin with the AI bot registry or the technical audit.

Technical

Open-source SEO agent tooling: what to build

The citation-collection harness, log-analysis pipeline and audit agents behind this site's own research are more credible open than closed.

8 min read · Sep 2026
Technical

Claude Code SEO failure modes

Where agentic SEO work actually goes wrong. Almost all existing coverage is promotional. This is the first-party failure catalogue instead.

17 min read · Sep 2026
Technical

Claude Code for SEO: the complete guide

Using Claude Code as an execution agent for technical SEO: audits, schema generation, structural fixes, log analysis. With a worked example and safety practices.

12 min read · Sep 2026
Technical

AI crawler statistics

How much of the web's bot traffic is AI. Which crawlers take the most relative to what they give back, and how that's shifted year over year. Verification status on every figure.

14 min read · Sep 2026
Technical

The AI Bot User-Agent Registry

Every AI crawler that matters for a website owner in one place: what it's for, how to verify it, and what to put in robots.txt. Maintained, not a one-off listicle.

12 min read · Sep 2026
Technical

GPTBot vs OAI-SearchBot: the AI crawler guide

The most-confused pair in AI SEO. One is a training crawler, one gets you cited. Plus every other AI bot and exactly what to allow or block.

11 min read · Jun 2026
Technical

The 50-point technical GEO audit

A complete technical checklist to make a site under 500 pages fully legible to Google, Bing and the AI answer engines.

17 min read · May 2026
Boundaries

What this programme will not publish

A research programme is defined as much by what it refuses to produce as by what it ships. These are standing commitments, not preferences.

No composite visibility score. Blending citation rate, mention volume and sentiment into one number produces something that moves for reasons nobody can explain. A score you cannot decompose is not a measurement, and it is the single most common product in this category.

No statistic without a traced origin. If a figure cannot be followed to a named study with a disclosed method, it does not get a number on these pages. It gets a sentence explaining where the chain broke.

No first-party claim without first-party data. Every experiment described here is labelled with whether it has run. Where a study is pre-registered and unrun, the page says so in the opening lines, not in a footnote.

No prediction presented as a finding. Forecasts about where AI search is heading are cheap and unfalsifiable. Where this site reasons ahead of the evidence, that reasoning is tiered as a hypothesis and can be checked later.

No quiet corrections. When a figure here turns out to be wrong, the page is corrected and the change is stated. Pre-registering predictions guarantees some will fail in public, which is the price of writing them down first.

Practical

Reading a statistic you find elsewhere

Most of the value of this programme is transferable. Five questions will tell you whether almost any AI-search figure means anything.

Who collected it, and what do they sell? Not a disqualifier. A vendor measuring the thing they sell access to has an interest in the direction of the answer, and you should know that before you quote it.

What was the query panel? A consumer panel and a professional panel produce different numbers from the same engine. The panel usually drives the result more than the engine does.

What is the denominator? A citation rate computed over queries where the engine never searched is diluted by definition. Most published rates never state the base.

Was variance measured? Ask the same engine the same question twice and the sources can change. Without a repeat run, part of any difference is noise being reported as a finding.

When was it collected? These products change without announcement. A figure with no date is not a statistic. It is a rumour with a decimal point.

The worked version of this is where AI SEO statistics come from, which traces the field's most-repeated figures back to their origins one at a time.

Legend

How to read the labels

Two separate systems run across these pages. They answer different questions and are not interchangeable.

LabelSystemWhat it tells you
Traceable Number provenanceThe figure leads to a named study with a disclosed method.
Partial Number provenanceA real source exists, but the method or sample is not fully disclosed.
Broken chain Number provenanceThe chain dead-ends. Someone is citing someone who cited nobody.
Fact Claim strengthDocumented and verifiable, usually from primary documentation.
Evidence Claim strengthSupported by published research, with the limits stated.
Hypothesis Claim strengthReasoning, not measurement. Plausible and unproven.
Open question Claim strengthNobody knows. Named so it can be answered later.

A grade describes where a number came from. A tier describes how much weight a claim can carry. A traceable number can still support only a hypothesis. The full method is in where AI SEO statistics come from.

Reuse

Using this research

Quote anything here, including the parts that are inconvenient. Two requests, both of which make the citation more useful to your own reader.

Carry the grade with the number. A figure graded partial is not the same as one graded traceable, and stripping the qualifier is how a soft number becomes a hard one across three reposts.

Carry the date. These surfaces change without announcement. An undated figure cannot be checked against anything, including itself a year later.

Cite this programme
Ritik Namdev (2026). AI Citation Index — pre-registration (v0). ritiknamdev.com/blog/ai-citation-index
When the data lands

Raw files ship with the first release, openly licensed and ungated. Nothing is available to download yet, and this page will say so until it is.

Read the dataset plan →
FAQ

Frequently asked questions

Are there finished studies with data on this site yet?
Not from the Citation Index. It is at v0, which is the pre-registration: the protocol, the metrics and the predictions, published before any data is collected. The first data release is scheduled for v1. Two pages do carry first-party log observations from this site’s own server, and they say so on the page.
Why publish a protocol before you have any results?
Because a method published afterwards can be quietly reshaped to fit whatever the data turned out to say. Committing the query set, the metrics and the predictions in advance is the cheapest protection against fooling ourselves, and it lets anyone check that the goalposts did not move.
Why are so many pages here about what is missing rather than what is known?
Because that is the honest state of the field. Most circulating AI-search statistics trace back to vendors whose data is their product, so the raw files are never published. Naming the gap is more useful than filling it with a number nobody can check.
Can I cite these pages?
Yes. Every page carries a version and a date, and every figure carries a grade showing how far we could trace it. Quote the grade along with the number. If a page says a figure is partial or broken-chain, that qualifier is part of the finding.
What happens if a study contradicts what this site already published?
The page gets corrected and the change is stated plainly. Pre-registering predictions means some of them will be wrong in public, which is the point of writing them down first.
How often does this hub change?
It is generated from the site’s content registry, so every research page that ships appears here automatically. The Citation Index status changes only when a version actually releases.
The Lab · Weekly

Get the first release first.

The Citation Index v1 goes to the newsletter before it is public, with the raw data attached. Join The Lab.

Free forever. Unsubscribe anytime.