The research question: which sources do AI search engines actually cite, how much does that vary between engines and between repeat runs of the same query, and does any of it hold still over time? Nobody can answer it from public evidence, because the vendors holding the largest citation corpora cannot release them — the data is the product. The Index is built to answer it in the open: a fixed query set, run across seven engines every quarter, every citation record published as a downloadable file. This page commits to the method in advance.
No data has been collected yet. Nothing on this page is a result. Version v0 is the pre-registration. The protocol, metrics and hypotheses are committed; collection of v1 begins Q1 2027.
- No open dataset of what AI search engines cite exists publicly — the largest citation corpora are held by vendors who cannot release them.
- Every figure quoted elsewhere on this site comes from third-party research and carries a provenance grade, not from this Index.
- Everything the Index is designed to measure: citation rates, variance across runs, and cross-engine concordance.
- Whether the seven-engine query set is large enough for per-vertical claims — v2 expands it for that reason.
Why does this dataset not already exist?
Pick any statistic circulating about AI search and try to trace it. Most of the time you land on a marketing blog citing another marketing blog, and the trail goes cold before it reaches a methodology. The field has an enormous amount of published conclusion and very little published evidence.
That is not because the research is bad — Ahrefs in particular has done the most rigorous public work in this space. The problem is structural: the organisations holding the largest citation datasets are visibility software companies, and a company whose product is proprietary citation data cannot open-source it without dismantling its own business. The incentive to publish a null result is close to zero.
The peer-reviewed exception, the Princeton-led GEO paper presented at KDD in 2024, ran a controlled 10,000-query benchmark and remains the only causal study most of the field can point to. Ahrefs' public re-measurement of its own AI Overview overlap figure — roughly 76% in mid-2025 down to roughly 38% by early 2026 — is the field's best example of a source correcting itself. What never appeared is an independent party whose entire purpose is measurement: a dataset anyone can download, re-run, and disagree with using their own numbers.
The AI search industry has an enormous amount of published conclusion and very little published evidence. The organisations with the best citation data are the ones who can least afford to release it.
Share on XWhat do we already know, and how well?
The current evidence base is thin and unevenly distributed. Source-mix figures exist for ChatGPT and Perplexity, are much weaker for Gemini, and are close to absent for Claude. Every figure this site republishes is graded in the statistics hub.
Three findings shape this design. Each is tagged with how much weight it can bear, and none has been independently replicated.
Google rank is a weak predictor of AI citation. Ahrefs found roughly 12% of URLs cited by ChatGPT, Gemini and Copilot rank in Google's top 10 for the same prompt; Perplexity is the outlier at closer to one in three. Ahrefs, 2026. Single-vendor, single time point.
Engines disagree sharply about sources. Analysis of a reported 680 million citations found ChatGPT skewing encyclopedic, Perplexity skewing Reddit and AI Overviews skewing multimodal — with only about 11% domain overlap between ChatGPT and Perplexity. Profound, 2026. Corpus not public; the figure cannot be independently checked.
Whether anything a site controls causally changes citation rate. Every published correlation is confounded — sites that do one thing well tend to do all of them well. No randomised, pre-registered intervention exists in public. No study located as of September 2026.
If ranking and citation were the same problem, all three bars would sit near 100%. They do not: whatever gets a page cited by ChatGPT is substantially not what gets it to rank, and the Google surfaces behave differently from the non-Google ones.
Each engine has a centre of gravity. If that holds, "AI visibility" is not one problem with one answer but at least three problems wearing the same name.
That is the economic backdrop to publisher hostility toward AI crawlers: a search engine taking roughly five pages per visitor it sends is a trade, one taking over two thousand is something else. Blocking a crawler is a visibility decision, which is why the robots.txt blocking census runs alongside this, and why what those crawlers can actually render is a prerequisite question.
What does the Index measure?
One fixed query set, run against seven surfaces, with every citation recorded. Not a score, not a composite index of visibility — a count of what appeared, where, and when.
| Surface | Retrieval | Why included |
|---|---|---|
| ChatGPT Search | Live retrieval via OAI-SearchBot | Largest consumer surface |
| Claude | Tool-invoked web search | Least-measured major surface |
| Gemini | Grounded generation | Conflated with Google's other surfaces everywhere |
| Perplexity | Own crawler and index | Most transparent — used to validate the pipeline |
| Google AI Overviews | Grounded over Google's index | Highest reach |
| Google AI Mode | Query fan-out, then synthesis | Where Google is heading |
| Bing Copilot | Bing index and grounding | Most tractable entry point for new sites |
Each citation record captures:
- URL, domain and domain category (publisher, brand, forum, encyclopedia, documentation, video)
- Position within the answer, and which sub-answer it supported
- The query, the engine, the run number and the timestamp
- The snippet or claim the citation was attached to
- Joined third-party metrics — domain rating, estimated traffic, domain age — recorded at collection time
What does a citation record look like?
A single illustrative record in the shape the Index will publish — invented values, showing structure, not a collected row:
{
"query_id": "q-0417",
"query_text": "best budget mirrorless camera 2027",
"engine": "chatgpt_search",
"run": 3,
"timestamp": "2027-01-14T09:32:11Z",
"citation": {
"url": "example.com/best-mirrorless-cameras",
"domain": "example.com",
"domain_category": "publisher",
"position": 2,
"supports_subanswer": "value_pick",
"snippet": "…the X-T50 remains the strongest sub-$1000 option…"
},
"domain_metrics": { "dr": 61, "est_traffic": 84000, "domain_age_years": 7 }
} Every field is joinable — DR against citation rate, domain age against citation position. That is the point of publishing raw records rather than a summary table: a summary answers the questions the publisher thought to ask, a raw file answers the ones nobody asked yet.
What is in the query set?
A study is only as good as what it asks, so the set is published in full — it is the part most likely to be criticised, which is exactly why it goes out in the open.
- Informational 400
- Commercial 250
- How-to 200
- Local 100
- Navigational 50
It is weighted toward informational and commercial intent, where AI answers displace clicks most visibly; navigational queries are a small slice because engines resolve them trivially. The ten verticals — consumer electronics, personal finance, health and wellness, home and garden, software and SaaS, travel, food and recipes, B2B services, education, local services — carry roughly 100 queries each at v1, enough for a directional per-vertical read only once v2 expands to 5,000.
The set does not change within a release cycle. Queries are added only at a version boundary, with the additions logged, so quarter-on-quarter figures stay comparable. Silently swapping queries between releases would make every trend line meaningless.
How is the data collected?
The whole point is that you can check this: the query set, the raw records and the collection code ship with every release.
- 01 Fixed query set Versioned, published, unchanged within a release
- 02 Five runs each Fresh session per run, no personalisation
- 03 Citation extraction URL, position, sub-answer, snippet, timestamp
- 04 Metric join DR, traffic, domain age recorded at collection time
- 05 Publish raw CSV + JSON + code + changelog, CC BY 4.0
Why 72 hours: a single instant risks catching a blip — a model rollback, a brief outage — and reporting it as the quarter's baseline, while a full month reintroduces the drift the window exists to control.
Why five runs, not one?
Variance is why single-run studies mislead. The same question can return a different source list minutes later, which is also the mechanism behind citation half-life and part of why query fan-out is hard to observe from outside.
Language models are not deterministic: ask the same question twice and you may get a different answer built from different sources. Almost every AI-visibility figure in circulation — including the ones charted above — comes from a single run per query. If run-to-run source disagreement is meaningful, some proportion of every published citation statistic is measurement error rather than signal, and nobody currently knows what that proportion is. Quantifying it is hypothesis H2.
Five runs does not eliminate the problem. It makes the problem visible. Every figure the Index publishes carries the spread across runs alongside the central estimate, so a reader can see whether a difference between two engines is real or within noise.
Almost every AI-visibility number in circulation comes from a single query run. If AI answers aren't deterministic, some share of every published statistic is measurement error — and nobody knows how much.
Share on XWhich metrics are reported?
Each of these exists because there is a real measurement problem with no current answer.
Citation Rate — the share of tracked queries in which a domain appears at least once. The base unit. Right now every vendor means something different by 'AI visibility'.
AI Share of Voice — a domain's citations as a share of all citations in a query set. Makes competitive comparison possible on a fixed denominator.
Citation Half-Life — days until a cited URL's citation rate falls to half its peak. Nobody has measured whether a citation persists at all.
Citation Efficiency — citations earned per 1,000 pages crawled. Connects server-log reality to visibility outcome. No equivalent metric exists.
One metric is deliberately not on that list: a composite "AI Visibility Score." Every visibility vendor has one, they are unfalsifiable, and building one here would undercut the only thing this project has going for it.
How is each metric calculated?
Each formula, then a worked example on invented numbers for a hypothetical domain, "acmegear.com."
| Metric | Formula |
|---|---|
| Citation Rate (CR) | queries citing the domain ÷ total tracked queries |
| Share of Voice (SOV) | domain's citations ÷ all citations across the query set |
| Citation Half-Life (CHL) | days elapsed when citation rate for a cohort of URLs first drops to 50% of its peak value |
| Citation Efficiency (CE) | (citations earned ÷ pages crawled by that engine's bot) × 1,000 |
acmegear.com: cited in 340 of the 1,000 tracked queries — a Citation Rate of 34%.
Same hypothetical: 340 citations out of 9,800 recorded that quarter gives a Share of Voice of roughly 3.5%, which accounts for how crowded the citation set was. If logs show an AI bot crawled 12,000 of its pages, Citation Efficiency is (340 ÷ 12,000) × 1,000 ≈ 28.3 citations per 1,000 pages crawled — a figure that only becomes interesting next to a competitor's.
Pre-registered hypotheses
These are posted now, before collection, with the direction predicted and the analysis specified. If the data contradicts them, that gets published as the result.
| # | Hypothesis | Predicted outcome | Current status |
|---|---|---|---|
| H1 | Cross-engine citation overlap is below 25% for the same query | Expect to hold | Correlational precedent only |
| H2 | Repeated identical queries return different source sets in over 30% of cases | Expect to hold | Untested in public |
| H3 | Google top-10 rank predicts citation for AI Overviews but not for ChatGPT or Claude | Expect to hold | Partial precedent (Ahrefs) |
| H4 | Citation rate is more concentrated than organic ranking — fewer domains take a larger share | Expect to hold | Unmeasured |
| H5 | A majority of cited URLs remain cited 90 days later | Expect not to hold | Never measured |
| H6 | Domain rating correlates with citation rate at r > 0.4 | Expect not to hold | Weak precedent suggests lower |
H5 and H6 are the interesting ones, because the prediction is that they fail. Registering a hypothesis you expect to reject is the cheapest available protection against reading a pattern into noise after the fact — and if H6 does hold, that is a genuinely surprising result worth more than a confirmation.
Limitations: what it will not prove
It will not establish causation, and it will not produce a single visibility score. Correlational work on mentions versus backlinks and the registered schema experiment exist precisely because observation alone cannot answer those questions.
The Index is an observational instrument, and observational data has hard limits that are easy to forget once a chart looks convincing.
- It cannot establish causation. If highly-cited domains share a trait, that does not mean the trait produced the citations. Causal claims need randomised intervention, which is a separate programme.
- It measures the engines, not your site. Citation rates on a fixed public query set describe the ecosystem. They do not tell an individual site owner what will happen to them.
- It is a snapshot of moving targets. These products change without notice. A finding is true of the surface as it behaved during a stated 72-hour window, and every figure is published with that window attached.
- Geography and language are fixed. Collection runs English-language queries from a single stated region. Multilingual and multi-region collection is planned but not funded, and no figure should be generalised beyond the language and region tested.
- Engine coverage stops at seven answer surfaces. Agentic browsers, MCP-mediated retrieval and in-app assistants are excluded because they do not emit a comparable citation list, so the Index describes AI answer citation, not every route an AI system takes to a page.
- Reproduction is only as good as the products allow. A third party re-running the published code in a later week is measuring a different surface state. Agreement on direction and magnitude is the realistic bar, not exact replication of a number.
- Sampling bias is real. A 1,000-query set is a choice, and a different set would produce different numbers. That is why the set is published rather than described.
- Residual personalisation may survive the controls. If an engine varies results by IP geography or device fingerprint even in a logged-out state, the recorded citation set may not match what every real user sees.
How can you check the work?
The collection code is intended for open release, described in open-source SEO agent tooling, and every planned file is listed in the dataset strategy. Results that contradict a prediction go to the null results registry.
An open study nobody audits is a closed study with extra steps. Three things ship with every release so the findings can be attacked. Any error found gets a dated changelog entry rather than a silent edit, credited to whoever found it.
When does each release land?
Quarterly full releases, monthly pulse updates in between. Each release is versioned and permanent: older versions stay online rather than being overwritten. If a release slips, the delay and the reason get posted here; if the project stops, that gets posted too.
- v0Sep 2026
Protocol only — this pre-registration.
Method committed before any collection.
- v1Q1 2027
1,000 queries × 7 engines × 5 runs.
Baseline citation rates, variance, cross-engine concordance.
- v2Q2 2027
Expands to 5,000 queries.
First per-vertical breakdowns.
- v3Q3 2027
5,000 queries, third collection window.
First citation half-life estimates become possible.
- v4Q4 2027
Full-year synthesis.
Twelve months of trend data and the annual report.
Get the first release, or contribute to it
There is no data yet, so the honest thing to offer is notice when there is: the newsletter carries each release and its raw files the day they publish. Two contributions would materially improve the work. Queries from your vertical that belong in the set, which is public and credits contributors. And anonymised server logs from a site you run, which unlock crawl-side measurements query data alone cannot reach. The about page explains how to get in touch.
To measure your own site against this standard before v1 lands: check which of your URLs are currently surfaced with the AI Overview Exposure Checker, then apply the definitions in the AI visibility measurement standard. To read the individual protocols this Index is sliced into: they are listed on the studies index.
Quick glossary
For every other term used across this site's research — GEO, AEO, query fan-out, crawl-to-referral ratio and more — see the full AI search glossary.
Namdev, R. (2026). The AI Citation Index (pre-registration) (v0). Retrieved from https://ritiknamdev.com/blog/ai-citation-index Published under CC BY 4.0 — reuse freely with attribution.
The measurement approach grew out of two first-party studies: a 90-day server-log test of llms.txt, which is where the crawl-side metrics come from, and Zero to Cited, which tracked a brand-new domain into AI search. For where the numbers on this page came from and how confident to be in each, see where AI SEO statistics actually come from.