Research & studies
Every research page on this site, in one place. The measurement programme is pre-registered and open: protocols before data, raw files alongside findings, and null results published like any other. This page says exactly what has been measured and what has not.
- The AI Citation Index is at v0 — pre-registration. The protocol is published. No Index data has been collected yet.
- 14 studies have their questions, methods and predictions published in advance, awaiting collection.
- 56 research pages are live, most of them tracing existing public figures to their origin and grading them.
- Two pages carry first-party observations from this site's own server logs. They are labelled as such.
- Where a number does not exist, these pages say so rather than estimating one.
Where the programme stands
The short version: the protocols are published and the data collection has not started. That is an unusual thing to advertise, so it is worth explaining why this page leads with it.
Most AI-search statistics in circulation come from companies whose data is the product they sell. That is not an accusation of bad faith. It is a structural problem. A vendor cannot publish the raw file without giving away the thing customers pay for, and they cannot publish a null result without undermining the category they sell into.
This programme occupies the position they cannot. Protocols are published before data is collected. Raw files ship alongside findings. A result that shows nothing gets published anyway. That last commitment is the one that costs something, and it is the clearest signal that a method is real.
The honest cost of that approach is visible right here. There is less finished data on this site than on a vendor blog, because the finished data has to survive being checked.
The AI Citation Index
The signature measurement: a recurring, versioned, openly published record of what AI search engines actually cite. Every other research page on this site orbits it.
- v0Sep 2026
Protocol only — the pre-registration.
Method, metrics and hypotheses committed before any collection.
- v1Q1 2027
1,000 queries × 7 engines × 5 runs.
Baseline citation rates, variance, cross-engine concordance.
Open question No citation data has been collected under this protocol yet. Any figure you see elsewhere on this site comes from third-party research and is graded for how far it could be traced. Read the full protocol in the AI Citation Index pre-registration, and the planned data releases in the dataset strategy.
How a study runs here
Four steps, in this order, without exception. The order is the whole point. Collecting first and deciding what counts afterwards is how an honest researcher accidentally produces a result that was never really there.
- 01 Pre-register Protocol, metrics and predictions published first
- 02 Collect Fixed query panel, disclosed conditions, repeated runs
- 03 Publish raw The dataset, not only the summary
- 04 Report either way A null result ships like any other
Step four is the one almost nobody commits to. Studies that find nothing are the ones that quietly disappear across every research field, which is why the published record looks more certain than the evidence is. Ours are listed in the null results registry before they run.
The collection code and analysis scripts are intended for open release, described in open-source SEO agent tooling. A finding you cannot reproduce is a claim, not a result.
Pre-registered studies
Each of these has its question, method and predictions published before collection. None has results yet. Each page states what it would take to run, and what a null result would look like.
| The question | Study | Status |
|---|---|---|
| Do different AI engines cite the same sources for the same question? | Cross-platform citation concordance | Open question |
| Do two Google surfaces cite the same sources for the same query? | AI Mode vs. AI Overviews: source delta | Open question |
| How long does a citation survive before it is replaced? | AI citation half-life: how long do citations last? | Open question |
| How long is the gap between a crawl and a first citation? | Crawl-to-citation latency | Open question |
| Does updating a page change how often it is cited? | Does content freshness increase AI citations? | Open question |
| Does structured data change citation rates? | Does schema markup increase AI citations? | Open question |
| Do visible author credentials affect citation? | Does an author bio affect AI citation? | Open question |
| How many sub-queries does one question actually become? | The Query Fan-Out Corpus | Open question |
| What gets cited when community discussion is thin? | The Reddit dependency: what Perplexity cites when Reddit has nothing | Open question |
| What replaces the encyclopaedia when it has no entry? | The Wikipedia dependency: what ChatGPT cites when Wikipedia has nothing | Open question |
| Do unlinked mentions predict citation better than links? | Brand mentions vs. backlinks: which predicts AI citation? | Open question |
| Which AI crawlers execute JavaScript? | Do AI crawlers render JavaScript? | Open question |
| What does an agentic browser actually request? | Agentic browsers: what do they actually fetch? | Open question |
| How many sites actually block AI crawlers? | The robots.txt AI-Blocking Census | Open question |
Statistics hubs
These pages take figures that already circulate and trace each one back to its origin. Where the chain dead-ends in a vendor blog citing another vendor blog, the page says so. That observation is usually worth more than the number.
The state of AI search: September 2026
A snapshot of where AI-search research actually stands as this site's research program launches. Five findings, five unanswered questions, and an honest inventory.
AI search CTR statistics
How click-through behavior changes when an AI Overview appears. The specific population each figure describes stays attached, not stripped away.
AI search and YMYL content: what is actually known
Health, finance, legal and safety queries carry higher stakes. Do AI engines cite differently for them? That is an open, under-studied question.
Local AI search: what is actually known
Local intent is a large share of all search behavior. Rigorous AI-search research on it is almost entirely absent. Stated plainly, not filled with a borrowed number.
AI shopping and agentic commerce statistics
AI agents are beginning to complete purchases on a shopper's behalf. What is actually measured, and how much of the coverage is still forecast rather than fact.
Agentic browsers: what do they actually fetch?
ChatGPT Atlas, Perplexity Comet and Gemini in Chrome are shipping now. What they actually request from a page has barely been studied. Pre-registered.
Do AI crawlers render JavaScript?
Retrieval bots may not execute client-side JavaScript. If not, any content that depends on it is invisible to them entirely. Pre-registered experiment.
AI Mode vs. AI Overviews: source delta
Two distinct Google AI surfaces, routinely treated as interchangeable. Do they actually cite the same sources for the same query? Pre-registered study.
The Reddit dependency: what Perplexity cites when Reddit has nothing
Perplexity's citations skew heavily to Reddit. For B2B and enterprise topics with thin discussion, what fills the gap? Pre-registered study.
The Wikipedia dependency: what ChatGPT cites when Wikipedia has nothing
ChatGPT's citations skew heavily to Wikipedia. For topics it doesn't cover well, what fills the gap? Pre-registered study.
Brand mentions vs. backlinks: which predicts AI citation?
Reported correlations put brand mentions at roughly 3x the predictive strength of backlinks for AI Overview visibility. A confound nobody has controlled for.
Does an author bio affect AI citation?
Author credentials and E-E-A-T signals are recommended constantly. No controlled test isolates whether they change AI citation rate. A pre-registered RCT.
Does content freshness increase AI citations?
'Keep content updated' is universal GEO advice. No controlled study has tested it against citation rate. A pre-registered RCT.
Crawl-to-citation latency
How many days pass between an AI crawler first fetching a new page and that page appearing in a citation? Pre-registered panel study.
The Query Fan-Out Corpus
Everyone repeats that Google AI Mode fans a query into '8-16 sub-queries.' Nobody has published the data behind that number. A pre-registered corpus design.
Cross-platform citation concordance
How much do ChatGPT, Perplexity, Gemini, Claude and Google agree on what to cite? A reported ~11% two-engine overlap suggests: not much. Pre-registered study.
AI citation half-life: how long do citations last?
Once a URL is cited by an AI engine, how long does it stay cited? Nobody has measured this. A 12-month tracking study, pre-registered.
Does schema markup increase AI citations?
The most contested tactic in GEO. This page registers a randomized test to settle it. Pre-registered before a single page is enrolled.
AI search conversion benchmarks
Published claims about AI traffic conversion range from 4.4x to 23x organic. Different orders of magnitude. Here is why, and the honest version of the statistic.
AI referral traffic statistics
How much traffic AI platforms actually send to websites, how it splits between engines, and why standard analytics likely undercount it.
Zero-click search statistics
The most commonly quoted zero-click figures, each attached to the specific population it actually measures. Collapsing them into one number is where most content goes wrong.
AI search market share statistics
How traffic is distributed across AI search platforms, how that's shifted over the past year, and how AI referral engagement compares to classic organic.
Gemini citation statistics
Gemini, AI Overviews, and AI Mode are three related but distinct Google surfaces, routinely conflated in coverage. What is actually known about Gemini specifically.
Perplexity citation statistics
Perplexity is the most transparent of the major AI search engines about its sources. Its citation pattern looks meaningfully different from ChatGPT.
Claude citation statistics
What is actually known about what Claude cites, and an honest accounting of how much isn't known. Claude is the least-studied major AI search surface.
ChatGPT citation statistics
What ChatGPT actually cites, how it retrieves information, and how weakly Google ranking predicts showing up in a ChatGPT Search answer.
Most-cited domains in AI search
Which kinds of sources dominate citations across ChatGPT, Perplexity and Google AI Overviews. And how little the major engines actually agree with each other.
Google AI Mode statistics and query fan-out, explained
How AI Mode actually retrieves information: the query fan-out mechanism, its documented patent basis, and what is confirmed versus estimated.
Google AI Overview statistics
Trigger rates, citation-to-ranking overlap, and click-through impact for AI Overviews - with a verification status attached to every figure.
The robots.txt AI-Blocking Census
How many of the web's most important sites block AI crawlers, measured directly from robots.txt, not quoted from a report with no visible sample. Pre-registration.
Where AI SEO statistics actually come from
We traced twelve of the most-quoted numbers in AI search back to their origin. Five hold up. Three dead-end. The most-repeated figure in the field is three studies wearing one number.
The AI Citation Index
A quarterly, open measurement of what seven AI search engines actually cite. Query set, raw data and methodology published in full. This is the pre-registration.
Zero to Cited: how a new site climbs into AI search
The evidence-backed timeline and playbook for taking a brand-new site from launch to its first citations in Bing, Perplexity, ChatGPT and AI Overviews.
Does llms.txt actually work? A 90-day log test
I deployed llms.txt and read the server logs for 90 days. Here is what the data, and the largest studies, actually say.
GEO research & guides
What the evidence supports, what it does not, and where the field is running ahead of its data. Start with what GEO actually is or the tactic evidence scoreboard.
The AI citation dataset strategy
Every dataset this site plans to publish, in one place: what it contains, when it ships, how it is licensed. Raw files, not just summary reports.
What is AI SEO?
The umbrella term for optimizing a site's visibility across the whole AI-search landscape: generative answers, AI crawlers, and the technical infrastructure behind both.
What is GEO?
Optimizing content so AI systems like ChatGPT, Perplexity and Gemini cite it when generating an answer. That is Generative Engine Optimization.
AI search glossary
Definitions for the terms used throughout this site's research. Kept tight, each linked to the fuller page where the concept is explored.
AI SEO statistics
Every AI-search statistic on this site, in one table, each with a verification grade and a link to full context.
AEO statistics
Answer Engine Optimization statistics. Genuinely fewer exist under this specific label than under GEO, and this page says so rather than padding itself.
GEO statistics
The most commonly cited Generative Engine Optimization statistics, each with a verification status attached rather than uniform, unearned confidence.
The Null-Results Registry
Every pre-registered prediction on this site that is, or is predicted to be, a null result, tracked in one place. Almost nobody publishes negative findings.
WebMCP and the agent-readable web
Google and Microsoft shipped an early preview of a structured agent-interaction protocol for websites in Feb 2026. What it is, and why site owners should care now.
MCP servers as a visibility channel
The Model Context Protocol lets AI agents query a business data directly. Does publishing one affect AI-search visibility? Untested, but genuinely early.
The AI Visibility Measurement Standard
'AI visibility' currently means whatever a vendor's dashboard measures. A proposed open specification so numbers from different sources become comparable.
GEO vs AEO vs LLMO: which term is winning
Three acronyms describe overlapping ideas about optimizing for AI-generated answers. The industry hasn't settled on one. What each means, where they overlap, and which is gaining ground.
The GEO Tactic Evidence Scoreboard
Every tactic the industry recommends for getting cited by AI search, graded by the actual evidence behind it. Not by how often it is repeated.
Bing Copilot SEO: the easiest AI engine to crack
ChatGPT search runs on Bing's index. Getting into Bing is the fastest path to your first AI citation. Here's the full playbook.
How to get cited by ChatGPT: 9 tactics, with data
The Princeton-backed levers - statistics, quotes, structure, freshness - that actually move AI citations.
Technical research
Crawlers, access controls, rendering and the plumbing that decides whether your pages are reachable at all. Begin with the AI bot registry or the technical audit.
Open-source SEO agent tooling: what to build
The citation-collection harness, log-analysis pipeline and audit agents behind this site's own research are more credible open than closed.
Claude Code SEO failure modes
Where agentic SEO work actually goes wrong. Almost all existing coverage is promotional. This is the first-party failure catalogue instead.
Claude Code for SEO: the complete guide
Using Claude Code as an execution agent for technical SEO: audits, schema generation, structural fixes, log analysis. With a worked example and safety practices.
AI crawler statistics
How much of the web's bot traffic is AI. Which crawlers take the most relative to what they give back, and how that's shifted year over year. Verification status on every figure.
The AI Bot User-Agent Registry
Every AI crawler that matters for a website owner in one place: what it's for, how to verify it, and what to put in robots.txt. Maintained, not a one-off listicle.
GPTBot vs OAI-SearchBot: the AI crawler guide
The most-confused pair in AI SEO. One is a training crawler, one gets you cited. Plus every other AI bot and exactly what to allow or block.
The 50-point technical GEO audit
A complete technical checklist to make a site under 500 pages fully legible to Google, Bing and the AI answer engines.
What this programme will not publish
A research programme is defined as much by what it refuses to produce as by what it ships. These are standing commitments, not preferences.
No composite visibility score. Blending citation rate, mention volume and sentiment into one number produces something that moves for reasons nobody can explain. A score you cannot decompose is not a measurement, and it is the single most common product in this category.
No statistic without a traced origin. If a figure cannot be followed to a named study with a disclosed method, it does not get a number on these pages. It gets a sentence explaining where the chain broke.
No first-party claim without first-party data. Every experiment described here is labelled with whether it has run. Where a study is pre-registered and unrun, the page says so in the opening lines, not in a footnote.
No prediction presented as a finding. Forecasts about where AI search is heading are cheap and unfalsifiable. Where this site reasons ahead of the evidence, that reasoning is tiered as a hypothesis and can be checked later.
No quiet corrections. When a figure here turns out to be wrong, the page is corrected and the change is stated. Pre-registering predictions guarantees some will fail in public, which is the price of writing them down first.
Reading a statistic you find elsewhere
Most of the value of this programme is transferable. Five questions will tell you whether almost any AI-search figure means anything.
Who collected it, and what do they sell? Not a disqualifier. A vendor measuring the thing they sell access to has an interest in the direction of the answer, and you should know that before you quote it.
What was the query panel? A consumer panel and a professional panel produce different numbers from the same engine. The panel usually drives the result more than the engine does.
What is the denominator? A citation rate computed over queries where the engine never searched is diluted by definition. Most published rates never state the base.
Was variance measured? Ask the same engine the same question twice and the sources can change. Without a repeat run, part of any difference is noise being reported as a finding.
When was it collected? These products change without announcement. A figure with no date is not a statistic. It is a rumour with a decimal point.
The worked version of this is where AI SEO statistics come from, which traces the field's most-repeated figures back to their origins one at a time.
How to read the labels
Two separate systems run across these pages. They answer different questions and are not interchangeable.
| Label | System | What it tells you |
|---|---|---|
| Traceable | Number provenance | The figure leads to a named study with a disclosed method. |
| Partial | Number provenance | A real source exists, but the method or sample is not fully disclosed. |
| Broken chain | Number provenance | The chain dead-ends. Someone is citing someone who cited nobody. |
| Fact | Claim strength | Documented and verifiable, usually from primary documentation. |
| Evidence | Claim strength | Supported by published research, with the limits stated. |
| Hypothesis | Claim strength | Reasoning, not measurement. Plausible and unproven. |
| Open question | Claim strength | Nobody knows. Named so it can be answered later. |
A grade describes where a number came from. A tier describes how much weight a claim can carry. A traceable number can still support only a hypothesis. The full method is in where AI SEO statistics come from.
Using this research
Quote anything here, including the parts that are inconvenient. Two requests, both of which make the citation more useful to your own reader.
Carry the grade with the number. A figure graded partial is not the same as one graded traceable, and stripping the qualifier is how a soft number becomes a hard one across three reposts.
Carry the date. These surfaces change without announcement. An undated figure cannot be checked against anything, including itself a year later.
Raw files ship with the first release, openly licensed and ungated. Nothing is available to download yet, and this page will say so until it is.
Read the dataset plan →Frequently asked questions
Are there finished studies with data on this site yet?
Why publish a protocol before you have any results?
Why are so many pages here about what is missing rather than what is known?
Can I cite these pages?
What happens if a study contradicts what this site already published?
How often does this hub change?
Get the first release first.
The Citation Index v1 goes to the newsletter before it is public, with the raw data attached. Join The Lab.