Original research · Pre-registered

The AI Citation Index

A quarterly, open measurement of what seven AI search engines actually cite — with the query set, the raw data and the collection code published in full.

This page is the pre-registration. It sets out the protocol, the metrics and the hypotheses before any data is collected, so that a null result is as publishable as a positive one.

Ritik Namdev Ritik Namdev ·Protocol published September 2026 ·24 min read ·Last verified September 2026 ·v0 — pre-registration ·v1 releases Q1 2027
The short version

The research question: which sources do AI search engines actually cite, how much does that vary between engines and between repeat runs of the same query, and does any of it hold still over time? Nobody can answer it from public evidence, because the vendors holding the largest citation corpora cannot release them — the data is the product. The Index is built to answer it in the open: a fixed query set, run across seven engines every quarter, every citation record published as a downloadable file. This page commits to the method in advance.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. Version v0 is the pre-registration. The protocol, metrics and hypotheses are committed; collection of v1 begins Q1 2027.

What is known
  • No open dataset of what AI search engines cite exists publicly — the largest citation corpora are held by vendors who cannot release them.
  • Every figure quoted elsewhere on this site comes from third-party research and carries a provenance grade, not from this Index.
What is not yet known
  • Everything the Index is designed to measure: citation rates, variance across runs, and cross-engine concordance.
  • Whether the seven-engine query set is large enough for per-vertical claims — v2 expands it for that reason.
Protocol summary Quotable in full. Every value below is stated elsewhere on this page; nothing here is a result.
Query count1,000 queries at v1, rising to 5,000 at v2.
EnginesSeven: ChatGPT Search, Claude, Gemini, Perplexity, Google AI Overviews, Google AI Mode, Bing Copilot.
Runs per queryFive, per engine, per collection window.
Collection windowA single 72-hour window per quarter; a monthly pulse re-runs a fixed 200-query subsample.
Primary outcomeCitation Rate — the share of tracked queries in which a domain is cited at least once. Supporting metrics: Share of Voice, Citation Half-Life, Citation Efficiency. No composite score.
Analysis planSix hypotheses (H1–H6) registered before collection, each with a predicted direction; variance across the five runs and confidence intervals reported alongside every figure; refusals and zero-citation answers recorded rather than dropped.
LicenceCC BY 4.0 — raw CSV and JSON, the query set and the collection code published with every release.
Release cadenceQuarterly full releases, monthly pulse updates in between. v0 (protocol) September 2026; v1 Q1 2027.

Why does this dataset not already exist?

Pick any statistic circulating about AI search and try to trace it. Most of the time you land on a marketing blog citing another marketing blog, and the trail goes cold before it reaches a methodology. The field has an enormous amount of published conclusion and very little published evidence.

That is not because the research is bad — Ahrefs in particular has done the most rigorous public work in this space. The problem is structural: the organisations holding the largest citation datasets are visibility software companies, and a company whose product is proprietary citation data cannot open-source it without dismantling its own business. The incentive to publish a null result is close to zero.

The peer-reviewed exception, the Princeton-led GEO paper presented at KDD in 2024, ran a controlled 10,000-query benchmark and remains the only causal study most of the field can point to. Ahrefs' public re-measurement of its own AI Overview overlap figure — roughly 76% in mid-2025 down to roughly 38% by early 2026 — is the field's best example of a source correcting itself. What never appeared is an independent party whose entire purpose is measurement: a dataset anyone can download, re-run, and disagree with using their own numbers.

The AI search industry has an enormous amount of published conclusion and very little published evidence. The organisations with the best citation data are the ones who can least afford to release it.

Share on X

What do we already know, and how well?

The current evidence base is thin and unevenly distributed. Source-mix figures exist for ChatGPT and Perplexity, are much weaker for Gemini, and are close to absent for Claude. Every figure this site republishes is graded in the statistics hub.

Three findings shape this design. Each is tagged with how much weight it can bear, and none has been independently replicated.

Evidence

Google rank is a weak predictor of AI citation. Ahrefs found roughly 12% of URLs cited by ChatGPT, Gemini and Copilot rank in Google's top 10 for the same prompt; Perplexity is the outlier at closer to one in three. Ahrefs, 2026. Single-vendor, single time point.

Evidence

Engines disagree sharply about sources. Analysis of a reported 680 million citations found ChatGPT skewing encyclopedic, Perplexity skewing Reddit and AI Overviews skewing multimodal — with only about 11% domain overlap between ChatGPT and Perplexity. Profound, 2026. Corpus not public; the figure cannot be independently checked.

Open question

Whether anything a site controls causally changes citation rate. Every published correlation is confounded — sites that do one thing well tend to do all of them well. No randomised, pre-registered intervention exists in public. No study located as of September 2026.

If ranking and citation were the same problem, all three bars would sit near 100%. They do not: whatever gets a page cited by ChatGPT is substantially not what gets it to rank, and the Google surfaces behave differently from the non-Google ones.

Ahrefs' two published measurements, nine months apart
Jul 2025Mar 2026
Two independent samples, not a tracked cohort: Ahrefs, Jul 2025 and Mar 2026. Nothing was measured in between, so no intermediate points are shown. Replacing this with an actually-measured quarterly series is one purpose of Index v1.

Each engine has a centre of gravity. If that holds, "AI visibility" is not one problem with one answer but at least three problems wearing the same name.

That is the economic backdrop to publisher hostility toward AI crawlers: a search engine taking roughly five pages per visitor it sends is a trade, one taking over two thousand is something else. Blocking a crawler is a visibility decision, which is why the robots.txt blocking census runs alongside this, and why what those crawlers can actually render is a prerequisite question.

What does the Index measure?

One fixed query set, run against seven surfaces, with every citation recorded. Not a score, not a composite index of visibility — a count of what appeared, where, and when.

Surfaces measured at v1
SurfaceRetrievalWhy included
ChatGPT SearchLive retrieval via OAI-SearchBotLargest consumer surface
ClaudeTool-invoked web searchLeast-measured major surface
GeminiGrounded generationConflated with Google's other surfaces everywhere
PerplexityOwn crawler and indexMost transparent — used to validate the pipeline
Google AI OverviewsGrounded over Google's indexHighest reach
Google AI ModeQuery fan-out, then synthesisWhere Google is heading
Bing CopilotBing index and groundingMost tractable entry point for new sites

Each citation record captures:

  • URL, domain and domain category (publisher, brand, forum, encyclopedia, documentation, video)
  • Position within the answer, and which sub-answer it supported
  • The query, the engine, the run number and the timestamp
  • The snippet or claim the citation was attached to
  • Joined third-party metrics — domain rating, estimated traffic, domain age — recorded at collection time

What does a citation record look like?

A single illustrative record in the shape the Index will publish — invented values, showing structure, not a collected row:

{
  "query_id": "q-0417",
  "query_text": "best budget mirrorless camera 2027",
  "engine": "chatgpt_search",
  "run": 3,
  "timestamp": "2027-01-14T09:32:11Z",
  "citation": {
    "url": "example.com/best-mirrorless-cameras",
    "domain": "example.com",
    "domain_category": "publisher",
    "position": 2,
    "supports_subanswer": "value_pick",
    "snippet": "…the X-T50 remains the strongest sub-$1000 option…"
  },
  "domain_metrics": { "dr": 61, "est_traffic": 84000, "domain_age_years": 7 }
}

Every field is joinable — DR against citation rate, domain age against citation position. That is the point of publishing raw records rather than a summary table: a summary answers the questions the publisher thought to ask, a raw file answers the ones nobody asked yet.

What is in the query set?

A study is only as good as what it asks, so the set is published in full — it is the part most likely to be criticised, which is exactly why it goes out in the open.

Composition of the v1 query set — 1,000 queries
  • Informational 400
  • Commercial 250
  • How-to 200
  • Local 100
  • Navigational 50
Stratified by intent class and spread across ten verticals. This is a design decision, not a finding: a different set would produce different numbers, which is the reason the file is published rather than described.

It is weighted toward informational and commercial intent, where AI answers displace clicks most visibly; navigational queries are a small slice because engines resolve them trivially. The ten verticals — consumer electronics, personal finance, health and wellness, home and garden, software and SaaS, travel, food and recipes, B2B services, education, local services — carry roughly 100 queries each at v1, enough for a directional per-vertical read only once v2 expands to 5,000.

The set does not change within a release cycle. Queries are added only at a version boundary, with the additions logged, so quarter-on-quarter figures stay comparable. Silently swapping queries between releases would make every trend line meaningless.

How is the data collected?

The whole point is that you can check this: the query set, the raw records and the collection code ship with every release.

Collection pipeline, per release
  1. 01 Fixed query set Versioned, published, unchanged within a release
  2. 02 Five runs each Fresh session per run, no personalisation
  3. 03 Citation extraction URL, position, sub-answer, snippet, timestamp
  4. 04 Metric join DR, traffic, domain age recorded at collection time
  5. 05 Publish raw CSV + JSON + code + changelog, CC BY 4.0
Query set 1,000 queries at v1, rising to 5,000 at v2. Stratified across five intent classes and ten verticals. Published as a versioned file so anyone can re-run it.
Repeat runs Every query runs five times per engine per collection window. Variance is reported alongside every figure rather than averaged away.
Collection window A single 72-hour window per quarter, to limit drift within a release. The monthly pulse re-runs a fixed 200-query subsample for the trend line.
Environment Fresh sessions, no personalisation, no logged-in history, consistent geography and language. Deviations are logged per run.
Exclusions Queries returning no answer, refusals, and answers with zero citations are recorded as such — not dropped. A refusal is data.
Published with each release Report page · methodology page · raw CSV and JSON · the query set · the collection code · a changelog · a permanent version identifier.

Why 72 hours: a single instant risks catching a blip — a model rollback, a brief outage — and reporting it as the quarter's baseline, while a full month reintroduces the drift the window exists to control.

Why five runs, not one?

Variance is why single-run studies mislead. The same question can return a different source list minutes later, which is also the mechanism behind citation half-life and part of why query fan-out is hard to observe from outside.

Language models are not deterministic: ask the same question twice and you may get a different answer built from different sources. Almost every AI-visibility figure in circulation — including the ones charted above — comes from a single run per query. If run-to-run source disagreement is meaningful, some proportion of every published citation statistic is measurement error rather than signal, and nobody currently knows what that proportion is. Quantifying it is hypothesis H2.

Five runs does not eliminate the problem. It makes the problem visible. Every figure the Index publishes carries the spread across runs alongside the central estimate, so a reader can see whether a difference between two engines is real or within noise.

Almost every AI-visibility number in circulation comes from a single query run. If AI answers aren't deterministic, some share of every published statistic is measurement error — and nobody knows how much.

Share on X

Which metrics are reported?

Each of these exists because there is a real measurement problem with no current answer.

CR

Citation Rate — the share of tracked queries in which a domain appears at least once. The base unit. Right now every vendor means something different by 'AI visibility'.

Defined here
SOV

AI Share of Voice — a domain's citations as a share of all citations in a query set. Makes competitive comparison possible on a fixed denominator.

Defined here
CHL

Citation Half-Life — days until a cited URL's citation rate falls to half its peak. Nobody has measured whether a citation persists at all.

Novel
CE

Citation Efficiency — citations earned per 1,000 pages crawled. Connects server-log reality to visibility outcome. No equivalent metric exists.

Novel

One metric is deliberately not on that list: a composite "AI Visibility Score." Every visibility vendor has one, they are unfalsifiable, and building one here would undercut the only thing this project has going for it.

How is each metric calculated?

Each formula, then a worked example on invented numbers for a hypothetical domain, "acmegear.com."

MetricFormula
Citation Rate (CR)queries citing the domain ÷ total tracked queries
Share of Voice (SOV)domain's citations ÷ all citations across the query set
Citation Half-Life (CHL)days elapsed when citation rate for a cohort of URLs first drops to 50% of its peak value
Citation Efficiency (CE)(citations earned ÷ pages crawled by that engine's bot) × 1,000
34%

acmegear.com: cited in 340 of the 1,000 tracked queries — a Citation Rate of 34%.

Worked example, invented figures

Same hypothetical: 340 citations out of 9,800 recorded that quarter gives a Share of Voice of roughly 3.5%, which accounts for how crowded the citation set was. If logs show an AI bot crawled 12,000 of its pages, Citation Efficiency is (340 ÷ 12,000) × 1,000 ≈ 28.3 citations per 1,000 pages crawled — a figure that only becomes interesting next to a competitor's.

Pre-registered hypotheses

These are posted now, before collection, with the direction predicted and the analysis specified. If the data contradicts them, that gets published as the result.

Hypotheses registered for v1, September 2026 — predictions made before collection, not results
#HypothesisPredicted outcomeCurrent status
H1Cross-engine citation overlap is below 25% for the same queryExpect to holdCorrelational precedent only
H2Repeated identical queries return different source sets in over 30% of casesExpect to holdUntested in public
H3Google top-10 rank predicts citation for AI Overviews but not for ChatGPT or ClaudeExpect to holdPartial precedent (Ahrefs)
H4Citation rate is more concentrated than organic ranking — fewer domains take a larger shareExpect to holdUnmeasured
H5A majority of cited URLs remain cited 90 days laterExpect not to holdNever measured
H6Domain rating correlates with citation rate at r > 0.4Expect not to holdWeak precedent suggests lower

H5 and H6 are the interesting ones, because the prediction is that they fail. Registering a hypothesis you expect to reject is the cheapest available protection against reading a pattern into noise after the fact — and if H6 does hold, that is a genuinely surprising result worth more than a confirmation.

Limitations: what it will not prove

It will not establish causation, and it will not produce a single visibility score. Correlational work on mentions versus backlinks and the registered schema experiment exist precisely because observation alone cannot answer those questions.

The Index is an observational instrument, and observational data has hard limits that are easy to forget once a chart looks convincing.

  • It cannot establish causation. If highly-cited domains share a trait, that does not mean the trait produced the citations. Causal claims need randomised intervention, which is a separate programme.
  • It measures the engines, not your site. Citation rates on a fixed public query set describe the ecosystem. They do not tell an individual site owner what will happen to them.
  • It is a snapshot of moving targets. These products change without notice. A finding is true of the surface as it behaved during a stated 72-hour window, and every figure is published with that window attached.
  • Geography and language are fixed. Collection runs English-language queries from a single stated region. Multilingual and multi-region collection is planned but not funded, and no figure should be generalised beyond the language and region tested.
  • Engine coverage stops at seven answer surfaces. Agentic browsers, MCP-mediated retrieval and in-app assistants are excluded because they do not emit a comparable citation list, so the Index describes AI answer citation, not every route an AI system takes to a page.
  • Reproduction is only as good as the products allow. A third party re-running the published code in a later week is measuring a different surface state. Agreement on direction and magnitude is the realistic bar, not exact replication of a number.
  • Sampling bias is real. A 1,000-query set is a choice, and a different set would produce different numbers. That is why the set is published rather than described.
  • Residual personalisation may survive the controls. If an engine varies results by IP geography or device fingerprint even in a logged-out state, the recorded citation set may not match what every real user sees.

How can you check the work?

The collection code is intended for open release, described in open-source SEO agent tooling, and every planned file is listed in the dataset strategy. Results that contradict a prediction go to the null results registry.

An open study nobody audits is a closed study with extra steps. Three things ship with every release so the findings can be attacked. Any error found gets a dated changelog entry rather than a silent edit, credited to whoever found it.

The query setIf the queries are badly chosen the results are wrong — and you can see the queries. Disagreement about the set is legitimate criticism, published alongside it.
The collection codeExtraction logic is where quiet errors live: a selector that misses one citation type skews everything downstream.
The raw recordsEvery individual citation row, not summary tables, so anyone can re-run the analysis and reach a different conclusion from the same data.

When does each release land?

Quarterly full releases, monthly pulse updates in between. Each release is versioned and permanent: older versions stay online rather than being overwritten. If a release slips, the delay and the reason get posted here; if the project stops, that gets posted too.

Planned releases
  1. v0Sep 2026

    Protocol only — this pre-registration.

    Method committed before any collection.

  2. v2Q2 2027

    Expands to 5,000 queries.

    First per-vertical breakdowns.

  3. v3Q3 2027

    5,000 queries, third collection window.

    First citation half-life estimates become possible.

  4. v4Q4 2027

    Full-year synthesis.

    Twelve months of trend data and the annual report.

Get the first release, or contribute to it

There is no data yet, so the honest thing to offer is notice when there is: the newsletter carries each release and its raw files the day they publish. Two contributions would materially improve the work. Queries from your vertical that belong in the set, which is public and credits contributors. And anonymised server logs from a site you run, which unlock crawl-side measurements query data alone cannot reach. The about page explains how to get in touch.

Where to go next

To measure your own site against this standard before v1 lands: check which of your URLs are currently surfaced with the AI Overview Exposure Checker, then apply the definitions in the AI visibility measurement standard. To read the individual protocols this Index is sliced into: they are listed on the studies index.

Quick glossary

CitationAn AI-generated answer referencing or linking to a specific source URL.
Citation Rate (CR)Share of tracked queries where a domain is cited at least once.
Share of Voice (SOV)A domain's citations as a share of all citations in the query set.
Citation Half-Life (CHL)Days until a URL's citation rate falls to half its peak.
Citation Efficiency (CE)Citations earned per 1,000 pages an AI bot crawled from the site.
Collection windowThe fixed period (72 hours per quarter) during which a release's data is gathered.

For every other term used across this site's research — GEO, AEO, query fan-out, crawl-to-referral ratio and more — see the full AI search glossary.

How to cite this
Namdev, R. (2026). The AI Citation Index (pre-registration) (v0). Retrieved from https://ritiknamdev.com/blog/ai-citation-index

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

The measurement approach grew out of two first-party studies: a 90-day server-log test of llms.txt, which is where the crawl-side metrics come from, and Zero to Cited, which tracked a brand-new domain into AI search. For where the numbers on this page came from and how confident to be in each, see where AI SEO statistics actually come from.

FAQ

Frequently asked questions

Why publish the methodology before collecting any data?
Because a hypothesis posted after the results are in is not a hypothesis — it is a description. Pre-registration is standard practice in science and almost entirely absent from SEO research. Posting the protocol first means a null result is as publishable as a positive one, which is the only way to keep the findings honest.
How is this different from Profound, Peec or the Semrush AI Toolkit?
Those are visibility products with proprietary corpora, and several of them are excellent. The difference is structural, not qualitative: their data is the product, so they cannot open-source it. The Index publishes the query set, the raw citation records and the collection code. You can re-run it and check the numbers.
Is a 1,000-query sample big enough?
For domain-level citation rates across seven engines, yes — the confidence intervals are reported alongside every figure. For long-tail or per-vertical questions it is not, which is why v2 expands to 5,000 queries. Any figure whose interval is too wide to support a claim will be published with that stated, not quietly dropped.
Why five runs per query instead of one?
Because AI answers are not deterministic. The same question can return different sources on a second attempt, so a single run measures noise as well as signal. Five runs allows the variance to be reported rather than hidden — and quantifying that variance is itself one of the first findings the Index will publish.
Can I use the data commercially?
Yes. Everything is CC BY 4.0 — reuse it, chart it, build on it, sell analysis of it. The only requirement is attribution back to the Index release you used, by version number.
What happens if the results contradict something you have already published?
The finding gets published anyway, along with a note about what changed and why. That includes contradicting my own earlier guides on this site.
Will the Index track paid placements or ads inside AI answers?
Not at v1. Sponsored placements inside AI answers are a newer, less standardised phenomenon across engines, and mixing them into an organic-citation dataset would muddy the metric before it even has a baseline. A separate tracker for sponsored placements is a candidate for a later version, tracked openly rather than folded in silently.
How do you handle an engine that changes its product mid-quarter?
The 72-hour collection window limits exposure to drift within a single release, but a major mid-quarter change (a new model version, a redesigned answer UI) gets logged in that release's changelog as a confound, not smoothed over. If the change is large enough to make the quarter's numbers unrepresentative, that gets stated plainly in the release notes rather than buried in a footnote.
Why not just use an existing academic benchmark instead of building a new query set?
Academic IR benchmarks are built for retrieval accuracy, not for citation-selection behaviour in commercial AI-search products, and none we reviewed cover all seven surfaces tracked here with a stratification aimed at SEO and GEO use cases specifically. Reusing the Princeton GEO paper's benchmark was considered, but at 10,000 queries built for the field as it stood in 2023-24, it predates AI Mode and Claude's web search tool entirely.
What stops this project from just becoming another vendor with a proprietary score?
Nothing except a standing commitment, which is why it's stated this explicitly: no composite score, raw data published every release, methodology first, and null results treated as findings rather than failures. If a future version of this page ever drops the raw-data commitment, that is the moment to stop trusting it — and the changelog will make any such change impossible to hide.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.