Nobody publishes open data on what AI search engines cite. The vendors who hold the largest citation corpora cannot release them, because the data is the product. The Index fills that gap: a fixed query set, run across seven engines every quarter, with every citation record published as a downloadable file. This page commits to the method in advance.
Why this exists
Pick any statistic circulating about AI search and try to trace it. Most of the time you land on a marketing blog citing another marketing blog, and the trail goes cold before it reaches a methodology. The field has an enormous amount of published conclusion and very little published evidence.
That is not because the research is bad. Some of it is very good — Ahrefs in particular has done the most rigorous public work in this space. The problem is structural. The organisations holding the largest citation datasets are visibility software companies, and a company whose product is proprietary citation data cannot open-source that data without dismantling its own business. The incentive to publish a null result is close to zero.
So the field is missing something specific: not more analysis, but an independent, reproducible, longitudinal measurement that anyone can check. That is the gap the Index is built for. It is not an attempt to out-market the visibility vendors — several of them do genuinely good analysis — it is an attempt to build the one thing none of them structurally can: a dataset that anyone can download, re-run, and disagree with using their own numbers rather than trusting the report.
The AI search industry has an enormous amount of published conclusion and very little published evidence. The organisations with the best citation data are the ones who can least afford to release it.
Share on XWho this is for
The Index is built to be useful to four different kinds of reader, each for a slightly different reason, and it is worth being explicit about that up front so the design decisions later in this page make sense in context.
If you fall into more than one of those categories — a practitioner who is also, say, building a competing tool — that is fine. The commitment to open data means there is no tier of access being withheld from anyone; a competitor gets the same raw files as a journalist.
How we got here: a short history of AI-search measurement
It is worth pausing on how the field arrived at its current, evidence-thin state, because the history explains the gap rather than just describing it. Generative AI answer products — ChatGPT, Perplexity, and later Google's AI Overviews and AI Mode — moved from novelty to mainstream search behaviour inside roughly two years. Academic research had a head start of sorts: the Princeton-led GEO paper, published in late 2023 and presented at KDD in 2024, ran a genuinely controlled 10,000-query benchmark testing specific content tactics against measurable visibility lift. It remains, as of this writing, the only peer-reviewed causal study most of the field can point to.
Commercial measurement moved faster but shallower. As soon as brands started asking "am I visible in ChatGPT," a wave of visibility-tracking startups appeared, each building a proprietary corpus of tracked prompts and citations to sell as a dashboard. That is a reasonable business, and some of these companies — Profound and Peec AI in particular — have built genuinely large corpora, reportedly running into the hundreds of millions of citations. But a corpus built to power a subscription product is, by construction, not a corpus built to be checked by a stranger. The methodology stays proprietary because disclosing it would let a competitor replicate the product.
Meanwhile, independent SEO publications — Ahrefs foremost among them — began publishing occasional, genuinely disclosed studies: sample sizes stated, methodology described, results updated when re-measured. Ahrefs' own re-measurement of its AI Overview overlap figure, from roughly 76% in mid-2025 down to roughly 38% by early 2026, is the best single example in the field of a source correcting itself in public rather than letting an old number keep circulating. That is the standard the Index is trying to generalise across every engine, every quarter, rather than treat as a happy exception.
What never appeared, in any of this, was an independent party whose entire purpose was measurement rather than either academic publication (slow, narrow in scope, not built for longitudinal tracking) or commercial dashboards (fast, broad, but closed). That is the specific, narrow gap this project is built to occupy — not to replace either of the existing categories, but to sit alongside them as the version that opens its books.
What we already know, and how well
Three findings shape the design of this study. Each is tagged with how much weight it can actually bear — and none of them has been independently replicated.
Google rank is a weak predictor of AI citation. Ahrefs found roughly 12% of URLs cited by ChatGPT, Gemini and Copilot rank in Google's top 10 for the same prompt; Perplexity is the outlier at closer to one in three. Ahrefs, 2026. Single-vendor, single time point.
Engines disagree sharply about sources. Analysis of a reported 680 million citations found ChatGPT skewing encyclopedic, Perplexity skewing Reddit and AI Overviews skewing multimodal — with only about 11% domain overlap between ChatGPT and Perplexity. Profound, 2026. Corpus not public; the figure cannot be independently checked.
Whether anything a site controls causally changes citation rate. Every published correlation is confounded — sites that do one thing well tend to do all of them well. No randomised, pre-registered intervention exists in public. No study located as of September 2026.
Read that chart carefully, because it is the single most consequential finding in the field. If ranking and citation were the same problem, all three bars would sit near 100%. They do not. Whatever gets a page cited by ChatGPT is substantially not the thing that gets it to rank — and the two Google surfaces behave differently from the two non-Google ones.
Each engine has a centre of gravity. ChatGPT leans encyclopedic, Perplexity leans toward forum discussion, AI Overviews lean multimodal. If that holds, "AI visibility" is not one problem with one answer — it is at least three different problems wearing the same name, and a strategy tuned for one engine may do nothing for another.
This is the economic backdrop, and it explains why publishers are increasingly hostile to AI crawlers. A search engine that takes roughly five pages per visitor it sends is a trade. One that takes over two thousand is something else. Any measurement of AI citation has to sit alongside this, because a site's decision to allow or block a crawler is a decision about visibility.
Common misconceptions this Index is designed to correct
Before describing the methodology, it is worth naming a handful of assumptions that circulate as settled fact in AI-search commentary but that the evidence above does not actually support.
"AI visibility" is a single, measurable thing. The source-skew data above suggests the opposite — three engines, three different retrieval logics, three different answers to "what gets cited." A single blended score across engines hides more than it reveals.
A high citation count today means a stable, durable citation. No public study has tracked the same citations forward in time to check whether they persist — which is exactly why Citation Half-Life is defined as a new metric below rather than assumed away.
Ranking well on Google is what gets you cited by AI. True for AI Overviews to a declining degree, weakly true for Perplexity, and only around 12% true for ChatGPT — a much weaker relationship than most GEO advice implies.
What the Index measures
One fixed query set, run against seven surfaces, with every citation recorded. Not a score, not a composite index of visibility — a count of what appeared, where, and when.
| Surface | Retrieval | Why included |
|---|---|---|
| ChatGPT Search | Live retrieval via OAI-SearchBot | Largest consumer surface |
| Claude | Tool-invoked web search | Least-measured major surface |
| Gemini | Grounded generation | Conflated with Google's other surfaces everywhere |
| Perplexity | Own crawler and index | Most transparent — used to validate the pipeline |
| Google AI Overviews | Grounded over Google's index | Highest reach |
| Google AI Mode | Query fan-out, then synthesis | Where Google is heading |
| Bing Copilot | Bing index and grounding | Most tractable entry point for new sites |
Each citation record captures:
- URL, domain and domain category (publisher, brand, forum, encyclopedia, documentation, video)
- Position within the answer, and which sub-answer it supported
- The query, the engine, the run number and the timestamp
- The snippet or claim the citation was attached to
- Joined third-party metrics — domain rating, estimated traffic, domain age — recorded at collection time
What a citation record actually looks like
Abstract field lists are hard to picture. Here is a single, illustrative record in the shape the Index will publish it — invented values, to show structure, not a real collected row:
{
"query_id": "q-0417",
"query_text": "best budget mirrorless camera 2027",
"engine": "chatgpt_search",
"run": 3,
"timestamp": "2027-01-14T09:32:11Z",
"citation": {
"url": "example.com/best-mirrorless-cameras",
"domain": "example.com",
"domain_category": "publisher",
"position": 2,
"supports_subanswer": "value_pick",
"snippet": "…the X-T50 remains the strongest sub-$1000 option…"
},
"domain_metrics": { "dr": 61, "est_traffic": 84000, "domain_age_years": 7 }
} Every field in that record is something a downstream analyst could join against other data — DR against citation rate, domain age against citation position, sub-answer type against domain category. That joinability is the actual point of publishing raw records instead of a summary table: a summary answers the questions the publisher thought to ask, a raw file answers the ones nobody has thought of yet.
The query set
A study is only as good as what it asks. Publishing the query set is what makes this checkable — and it is the part most likely to be criticised, which is exactly why it goes out in the open.
- Informational 400
- Commercial 250
- How-to 200
- Local 100
- Navigational 50
The set is weighted toward informational and commercial intent because that is where AI answers are displacing clicks most visibly. Navigational queries are deliberately a small slice — engines resolve them trivially and they tell you little about source selection. Local queries are included at a modest weight because local intent is almost entirely absent from published AI-search research, and it is worth knowing whether the patterns hold there at all.
Once published, the set does not change within a release cycle. Queries can be added at a version boundary, with the additions logged, so that quarter-on-quarter figures stay comparable. Silently swapping queries between releases would make every trend line meaningless.
The ten verticals spread across the set are chosen for a mix of commercial relevance and research interest rather than pure representativeness of "the web": consumer electronics, personal finance, health and wellness, home and garden, software and SaaS, travel, food and recipes, B2B services, education, and local services. Each vertical carries roughly 100 queries at v1, enough for a directional per-vertical read once v2 expands the total set to 5,000 and pushes each vertical's sample past the point where a single anomalous query can swing the whole category's number.
Methodology
The whole point is that you can check this. The query set, the raw records and the collection code are published with every release.
- 01 Fixed query set Versioned, published, unchanged within a release
- 02 Five runs each Fresh session per run, no personalisation
- 03 Citation extraction URL, position, sub-answer, snippet, timestamp
- 04 Metric join DR, traffic, domain age recorded at collection time
- 05 Publish raw CSV + JSON + code + changelog, CC BY 4.0
One methodological choice deserves its own explanation: why 72 hours, specifically, rather than a single instant or a full month. A single instant risks catching a temporary blip — a model rollback, a brief outage, an unrelated product incident — and reporting it as the quarter's baseline. A full month reintroduces the drift problem the window is meant to control for, since a product can change meaningfully within thirty days. Seventy-two hours is a compromise: long enough to average out a single anomalous hour, short enough that "this is what the surface looked like in this window" stays a meaningful, falsifiable statement.
The variance problem
This deserves its own section, because it undermines a great deal of what is currently published as fact.
Language models are not deterministic. Ask the same question twice and you may get a different answer built from different sources. Almost every AI-visibility figure in circulation — including the ones charted above — comes from a single run per query. If the re-run rate of source disagreement is meaningful, then some proportion of every published citation statistic is measurement error rather than signal, and nobody currently knows what that proportion is.
Five runs per query does not eliminate the problem. It makes the problem visible. Every figure the Index publishes will carry the spread across runs alongside the central estimate, so a reader can see whether a difference between two engines is real or within noise. Quantifying that spread is hypothesis H2, and it is the first thing v1 will report.
Almost every AI-visibility number in circulation comes from a single query run. If AI answers aren't deterministic, some share of every published statistic is measurement error — and nobody knows how much.
Share on XConsider a concrete, hypothetical illustration of why this matters practically. Suppose Domain A shows a 40% citation rate on a single-run measurement and Domain B shows 34%. A report built on that single run would confidently declare Domain A "more visible." But if five runs on Domain A actually range from 28% to 52%, and five runs on Domain B range from 30% to 38%, the honest conclusion is that the two domains are statistically indistinguishable — the single-run comparison manufactured a difference that the repeated measurement dissolves. This is not a hypothetical failure mode unique to this example; it is the default failure mode of every single-run citation study published to date, including several referenced elsewhere on this site.
Four metrics worth defining
New terminology should earn its place. Each of these exists because there is a real measurement problem with no current answer — not to coin a phrase.
Citation Rate — the share of tracked queries in which a domain appears at least once. The base unit. Right now every vendor means something different by 'AI visibility'.
AI Share of Voice — a domain's citations as a share of all citations in a query set. Makes competitive comparison possible on a fixed denominator.
Citation Half-Life — days until a cited URL's citation rate falls to half its peak. Nobody has measured whether a citation persists at all.
Citation Efficiency — citations earned per 1,000 pages crawled. Connects server-log reality to visibility outcome. No equivalent metric exists.
There is one metric deliberately not on that list: a composite "AI Visibility Score". Every visibility vendor has one, they are unfalsifiable, and building one here would undercut the only thing this project has going for it.
How the metrics are calculated, with a worked example
Definitions are easier to trust when the arithmetic behind them is visible. Below is each formula, followed by a single worked example using invented numbers for a hypothetical domain, "acmegear.com."
| Metric | Formula |
|---|---|
| Citation Rate (CR) | queries citing the domain ÷ total tracked queries |
| Share of Voice (SOV) | domain's citations ÷ all citations across the query set |
| Citation Half-Life (CHL) | days elapsed when citation rate for a cohort of URLs first drops to 50% of its peak value |
| Citation Efficiency (CE) | (citations earned ÷ pages crawled by that engine's bot) × 1,000 |
acmegear.com: cited in 340 of the 1,000 tracked queries — a Citation Rate of 34%.
Continuing the same hypothetical: if acmegear.com earned 340 citations out of 9,800 total citations recorded across the whole query set that quarter, its Share of Voice would be 340 ÷ 9,800, or roughly 3.5% — a meaningfully different, and arguably more useful, number than the bare citation rate, because it accounts for how crowded the overall citation landscape was that quarter. If server logs show an AI bot crawled 12,000 of acmegear.com's pages that same quarter, its Citation Efficiency would be (340 ÷ 12,000) × 1,000 ≈ 28.3 citations per 1,000 pages crawled — a number that becomes genuinely interesting only once compared against a competitor's efficiency figure, since a site that earns the same citation count from a tenth of the crawl volume is doing something structurally different with its content.
Pre-registered hypotheses
These are posted now, before collection, with the direction predicted and the analysis specified. If the data contradicts them, that gets published as the result.
| # | Hypothesis | Prediction | Current status |
|---|---|---|---|
| H1 | Cross-engine citation overlap is below 25% for the same query | Supported | Correlational precedent only |
| H2 | Repeated identical queries return different source sets in over 30% of cases | Supported | Untested in public |
| H3 | Google top-10 rank predicts citation for AI Overviews but not for ChatGPT or Claude | Supported | Partial precedent (Ahrefs) |
| H4 | Citation rate is more concentrated than organic ranking — fewer domains take a larger share | Supported | Unmeasured |
| H5 | A majority of cited URLs remain cited 90 days later | Not supported | Never measured |
| H6 | Domain rating correlates with citation rate at r > 0.4 | Not supported | Weak precedent suggests lower |
H5 and H6 are the interesting ones, because the prediction is that they fail. Registering a hypothesis you expect to reject is the cheapest available protection against reading a pattern into noise after the fact — and if H6 does hold, that is a genuinely surprising result worth more than a confirmation.
How this compares to existing visibility tools
The Index is not trying to replace a paid visibility dashboard for day-to-day monitoring — those tools are built for a different job (continuous tracking of your own brand across many prompts) and do it well. The comparison below is about what each category is structurally able to offer, not about which product is "better."
| Property | Commercial visibility tools | The AI Citation Index |
|---|---|---|
| Raw data downloadable | No — dashboard only | Yes, every release |
| Methodology fully disclosed | Partial, varies by vendor | Published before collection |
| Continuous monitoring of your own brand | Yes — this is their core product | No — fixed public query set only |
| Cost to access | Subscription | Free, CC BY 4.0 |
| Multi-run variance reported | Rarely disclosed | Always, per release |
| Reproducible by a third party | No | Yes — code and query set are public |
In practice, the two are complementary: a brand that wants to track its own daily citation rate across a custom prompt set still needs a commercial tool. A researcher, journalist, or another tool builder who needs a checkable, reproducible baseline for the field as a whole has, as of this writing, nowhere else to go — which is the specific need this project is built to serve.
What the Index will not prove
This matters as much as the design. The Index is an observational instrument, and observational data has hard limits that are easy to forget once a chart looks convincing.
- It cannot establish causation. If highly-cited domains share a trait, that does not mean the trait produced the citations. Causal claims need randomised intervention, which is a separate programme.
- It measures the engines, not your site. Citation rates on a fixed public query set describe the ecosystem. They do not tell an individual site owner what will happen to them.
- It is a snapshot of moving targets. These products change without notice. A finding is true of the surface as it behaved during a stated 72-hour window, and every figure is published with that window attached.
- It is English-first at v1. Multilingual collection is planned but not funded yet, and results should not be generalised beyond the language tested.
- Sampling bias is real. A 1,000-query set is a choice, and a different set would produce different numbers. That is exactly why the set is published rather than described.
- It cannot detect a citation an engine never surfaces to the collection environment specifically. If an engine personalises results even in a logged-out state based on IP geography or device fingerprint in ways the collection setup does not fully neutralise, the recorded citation set may not match what every real user sees.
How to check our work
An open study that nobody audits is just a closed study with extra steps. Three things are published specifically so the findings can be attacked:
- The query set. If the queries are badly chosen, the results are wrong, and you can see the queries. Disagreement about the set is legitimate criticism and gets published alongside it.
- The collection code. Extraction logic is where quiet errors live — a selector that misses a citation type will skew everything downstream. The code ships with each release.
- The raw records. Not summary tables. Every individual citation row, so anyone can re-run the analysis and reach a different conclusion from the same data.
If you find an error, it gets corrected in a dated changelog entry rather than silently edited — and the person who found it gets credited.
Release schedule
Quarterly full releases, monthly pulse updates in between. Each release is versioned and permanent — older versions stay online rather than being overwritten, so historical figures remain checkable.
- v0Sep 2026
Protocol only — this pre-registration.
Method committed before any collection.
- v1Q1 2027
1,000 queries × 7 engines × 5 runs.
Baseline citation rates, variance, cross-engine concordance.
- v2Q2 2027
Expands to 5,000 queries.
First per-vertical breakdowns.
- v3Q3 2027
5,000 queries, third collection window.
First citation half-life estimates become possible.
- v4Q4 2027
Full-year synthesis.
Twelve months of trend data and the annual report.
If a release slips, the delay gets posted here with the reason. If the project stops, that gets posted too. An abandoned research programme that quietly goes stale is worse than one that never started.
Contribute a query, or your logs
Two things would materially improve this work. If you have queries in your vertical that you think belong in the set, send them — the set is public and contributions are credited. And if you run a site and would share anonymised server logs, that unlocks the crawl-side measurements that the query data alone cannot reach. Both routes are open at hello@ritiknamdev.com.
Quick glossary
For every other term used across this site's research — GEO, AEO, query fan-out, crawl-to-referral ratio and more — see the full AI search glossary.
Namdev, R. (2026). The AI Citation Index (pre-registration) (v0). Retrieved from https://ritiknamdev.com/blog/ai-citation-index Published under CC BY 4.0 — reuse freely with attribution.
The measurement approach here grew out of two first-party studies: a 90-day server-log test of llms.txt, which is where the crawl-side metrics come from, and Zero to Cited, which tracked a brand-new domain into AI search. For the technical groundwork, see GPTBot vs OAI-SearchBot and the 50-point technical GEO audit. For where the numbers on this page came from and how confident to be in each, see where AI SEO statistics actually come from.