Original research · Pre-registered

The AI Citation Index

A quarterly, open measurement of what seven AI search engines actually cite — with the query set, the raw data and the collection code published in full.

This page is the pre-registration. It sets out the protocol, the metrics and the hypotheses before any data is collected, so that a null result is as publishable as a positive one.

Ritik Namdev Ritik Namdev ·Protocol published September 2026 ·v0 — pre-registration ·v1 releases Q1 2027
The short version

Nobody publishes open data on what AI search engines cite. The vendors who hold the largest citation corpora cannot release them, because the data is the product. The Index fills that gap: a fixed query set, run across seven engines every quarter, with every citation record published as a downloadable file. This page commits to the method in advance.

Why this exists

Pick any statistic circulating about AI search and try to trace it. Most of the time you land on a marketing blog citing another marketing blog, and the trail goes cold before it reaches a methodology. The field has an enormous amount of published conclusion and very little published evidence.

That is not because the research is bad. Some of it is very good — Ahrefs in particular has done the most rigorous public work in this space. The problem is structural. The organisations holding the largest citation datasets are visibility software companies, and a company whose product is proprietary citation data cannot open-source that data without dismantling its own business. The incentive to publish a null result is close to zero.

So the field is missing something specific: not more analysis, but an independent, reproducible, longitudinal measurement that anyone can check. That is the gap the Index is built for. It is not an attempt to out-market the visibility vendors — several of them do genuinely good analysis — it is an attempt to build the one thing none of them structurally can: a dataset that anyone can download, re-run, and disagree with using their own numbers rather than trusting the report.

The AI search industry has an enormous amount of published conclusion and very little published evidence. The organisations with the best citation data are the ones who can least afford to release it.

Share on X

Who this is for

The Index is built to be useful to four different kinds of reader, each for a slightly different reason, and it is worth being explicit about that up front so the design decisions later in this page make sense in context.

SEO & GEO practitionersWant a citable, checkable number instead of a vendor bullet point.
Journalists & researchersNeed a source that discloses method, not just a headline stat.
Site owners deciding on AI crawlersWant real data before blocking or allowing a bot.
Other tool buildersCan build on the raw data instead of re-collecting it from scratch.

If you fall into more than one of those categories — a practitioner who is also, say, building a competing tool — that is fine. The commitment to open data means there is no tier of access being withheld from anyone; a competitor gets the same raw files as a journalist.

How we got here: a short history of AI-search measurement

It is worth pausing on how the field arrived at its current, evidence-thin state, because the history explains the gap rather than just describing it. Generative AI answer products — ChatGPT, Perplexity, and later Google's AI Overviews and AI Mode — moved from novelty to mainstream search behaviour inside roughly two years. Academic research had a head start of sorts: the Princeton-led GEO paper, published in late 2023 and presented at KDD in 2024, ran a genuinely controlled 10,000-query benchmark testing specific content tactics against measurable visibility lift. It remains, as of this writing, the only peer-reviewed causal study most of the field can point to.

Commercial measurement moved faster but shallower. As soon as brands started asking "am I visible in ChatGPT," a wave of visibility-tracking startups appeared, each building a proprietary corpus of tracked prompts and citations to sell as a dashboard. That is a reasonable business, and some of these companies — Profound and Peec AI in particular — have built genuinely large corpora, reportedly running into the hundreds of millions of citations. But a corpus built to power a subscription product is, by construction, not a corpus built to be checked by a stranger. The methodology stays proprietary because disclosing it would let a competitor replicate the product.

Meanwhile, independent SEO publications — Ahrefs foremost among them — began publishing occasional, genuinely disclosed studies: sample sizes stated, methodology described, results updated when re-measured. Ahrefs' own re-measurement of its AI Overview overlap figure, from roughly 76% in mid-2025 down to roughly 38% by early 2026, is the best single example in the field of a source correcting itself in public rather than letting an old number keep circulating. That is the standard the Index is trying to generalise across every engine, every quarter, rather than treat as a happy exception.

What never appeared, in any of this, was an independent party whose entire purpose was measurement rather than either academic publication (slow, narrow in scope, not built for longitudinal tracking) or commercial dashboards (fast, broad, but closed). That is the specific, narrow gap this project is built to occupy — not to replace either of the existing categories, but to sit alongside them as the version that opens its books.

What we already know, and how well

Three findings shape the design of this study. Each is tagged with how much weight it can actually bear — and none of them has been independently replicated.

Evidence

Google rank is a weak predictor of AI citation. Ahrefs found roughly 12% of URLs cited by ChatGPT, Gemini and Copilot rank in Google's top 10 for the same prompt; Perplexity is the outlier at closer to one in three. Ahrefs, 2026. Single-vendor, single time point.

Evidence

Engines disagree sharply about sources. Analysis of a reported 680 million citations found ChatGPT skewing encyclopedic, Perplexity skewing Reddit and AI Overviews skewing multimodal — with only about 11% domain overlap between ChatGPT and Perplexity. Profound, 2026. Corpus not public; the figure cannot be independently checked.

Open question

Whether anything a site controls causally changes citation rate. Every published correlation is confounded — sites that do one thing well tend to do all of them well. No randomised, pre-registered intervention exists in public. No study located as of September 2026.

Share of AI-cited URLs that also rank in Google's top 10
Perplexity
~33%
AI Overviews
~38%
ChatGPT / Gemini / Copilot
~12%
Source: Ahrefs, 2026, for the same prompt. The AI Overviews figure is reported as declining year over year — one of the first things v1 will attempt to replicate independently.

Read that chart carefully, because it is the single most consequential finding in the field. If ranking and citation were the same problem, all three bars would sit near 100%. They do not. Whatever gets a page cited by ChatGPT is substantially not the thing that gets it to rank — and the two Google surfaces behave differently from the two non-Google ones.

Illustrative shape of the AIO / top-10 overlap decline (not new data)
Q3 '25Q4 '25Q1 '26Q2 '26Q1 '26 (Ahrefs)
This line connects Ahrefs' two published data points (mid-2025 and early 2026) with an illustrative, evenly-spaced curve to show the shape of the decline — the intermediate quarters are not independently measured and should not be read as real data points. One purpose of Index v1 is to replace this illustration with an actually-measured quarterly line.
Dominant source type, as a share of each engine's top citations
ChatGPT → Wikipedia
47.9%
Perplexity → Reddit
46.7%
AI Overviews → YouTube
23.3%
Source: Profound / Discovered Labs, 2026, from a reported 680M-citation corpus. The underlying data is not public, so these figures are reproduced here as claims rather than verified findings.

Each engine has a centre of gravity. ChatGPT leans encyclopedic, Perplexity leans toward forum discussion, AI Overviews lean multimodal. If that holds, "AI visibility" is not one problem with one answer — it is at least three different problems wearing the same name, and a strategy tuned for one engine may do nothing for another.

Pages crawled per referral sent back — log scale
Mistral
3,389 : 1
Anthropic
2,237 : 1
OpenAI
217 : 1
Google
4.6 : 1
Source: Cloudflare-derived figures, July 2026, via secondary aggregators. Bars are log₁₀ scaled because the raw range spans three orders of magnitude. Google's traditional crawler is included for contrast.

This is the economic backdrop, and it explains why publishers are increasingly hostile to AI crawlers. A search engine that takes roughly five pages per visitor it sends is a trade. One that takes over two thousand is something else. Any measurement of AI citation has to sit alongside this, because a site's decision to allow or block a crawler is a decision about visibility.

Common misconceptions this Index is designed to correct

Before describing the methodology, it is worth naming a handful of assumptions that circulate as settled fact in AI-search commentary but that the evidence above does not actually support.

Hypothesis

"AI visibility" is a single, measurable thing. The source-skew data above suggests the opposite — three engines, three different retrieval logics, three different answers to "what gets cited." A single blended score across engines hides more than it reveals.

Hypothesis

A high citation count today means a stable, durable citation. No public study has tracked the same citations forward in time to check whether they persist — which is exactly why Citation Half-Life is defined as a new metric below rather than assumed away.

Hypothesis

Ranking well on Google is what gets you cited by AI. True for AI Overviews to a declining degree, weakly true for Perplexity, and only around 12% true for ChatGPT — a much weaker relationship than most GEO advice implies.

What the Index measures

One fixed query set, run against seven surfaces, with every citation recorded. Not a score, not a composite index of visibility — a count of what appeared, where, and when.

Surfaces measured at v1
SurfaceRetrievalWhy included
ChatGPT SearchLive retrieval via OAI-SearchBotLargest consumer surface
ClaudeTool-invoked web searchLeast-measured major surface
GeminiGrounded generationConflated with Google's other surfaces everywhere
PerplexityOwn crawler and indexMost transparent — used to validate the pipeline
Google AI OverviewsGrounded over Google's indexHighest reach
Google AI ModeQuery fan-out, then synthesisWhere Google is heading
Bing CopilotBing index and groundingMost tractable entry point for new sites

Each citation record captures:

  • URL, domain and domain category (publisher, brand, forum, encyclopedia, documentation, video)
  • Position within the answer, and which sub-answer it supported
  • The query, the engine, the run number and the timestamp
  • The snippet or claim the citation was attached to
  • Joined third-party metrics — domain rating, estimated traffic, domain age — recorded at collection time

What a citation record actually looks like

Abstract field lists are hard to picture. Here is a single, illustrative record in the shape the Index will publish it — invented values, to show structure, not a real collected row:

{
  "query_id": "q-0417",
  "query_text": "best budget mirrorless camera 2027",
  "engine": "chatgpt_search",
  "run": 3,
  "timestamp": "2027-01-14T09:32:11Z",
  "citation": {
    "url": "example.com/best-mirrorless-cameras",
    "domain": "example.com",
    "domain_category": "publisher",
    "position": 2,
    "supports_subanswer": "value_pick",
    "snippet": "…the X-T50 remains the strongest sub-$1000 option…"
  },
  "domain_metrics": { "dr": 61, "est_traffic": 84000, "domain_age_years": 7 }
}

Every field in that record is something a downstream analyst could join against other data — DR against citation rate, domain age against citation position, sub-answer type against domain category. That joinability is the actual point of publishing raw records instead of a summary table: a summary answers the questions the publisher thought to ask, a raw file answers the ones nobody has thought of yet.

The query set

A study is only as good as what it asks. Publishing the query set is what makes this checkable — and it is the part most likely to be criticised, which is exactly why it goes out in the open.

Composition of the v1 query set — 1,000 queries
  • Informational 400
  • Commercial 250
  • How-to 200
  • Local 100
  • Navigational 50
Stratified by intent class and spread across ten verticals. This is a design decision, not a finding: a different set would produce different numbers, which is the reason the file is published rather than described.

The set is weighted toward informational and commercial intent because that is where AI answers are displacing clicks most visibly. Navigational queries are deliberately a small slice — engines resolve them trivially and they tell you little about source selection. Local queries are included at a modest weight because local intent is almost entirely absent from published AI-search research, and it is worth knowing whether the patterns hold there at all.

Once published, the set does not change within a release cycle. Queries can be added at a version boundary, with the additions logged, so that quarter-on-quarter figures stay comparable. Silently swapping queries between releases would make every trend line meaningless.

The ten verticals spread across the set are chosen for a mix of commercial relevance and research interest rather than pure representativeness of "the web": consumer electronics, personal finance, health and wellness, home and garden, software and SaaS, travel, food and recipes, B2B services, education, and local services. Each vertical carries roughly 100 queries at v1, enough for a directional per-vertical read once v2 expands the total set to 5,000 and pushes each vertical's sample past the point where a single anomalous query can swing the whole category's number.

Methodology

The whole point is that you can check this. The query set, the raw records and the collection code are published with every release.

Collection pipeline, per release
  1. 01 Fixed query set Versioned, published, unchanged within a release
  2. 02 Five runs each Fresh session per run, no personalisation
  3. 03 Citation extraction URL, position, sub-answer, snippet, timestamp
  4. 04 Metric join DR, traffic, domain age recorded at collection time
  5. 05 Publish raw CSV + JSON + code + changelog, CC BY 4.0
Query set 1,000 queries at v1, rising to 5,000 at v2. Stratified across five intent classes and ten verticals. Published as a versioned file so anyone can re-run it.
Repeat runs Every query runs five times per engine per collection window. Variance is reported alongside every figure rather than averaged away.
Collection window A single 72-hour window per quarter, to limit drift within a release. The monthly pulse re-runs a fixed 200-query subsample for the trend line.
Environment Fresh sessions, no personalisation, no logged-in history, consistent geography and language. Deviations are logged per run.
Exclusions Queries returning no answer, refusals, and answers with zero citations are recorded as such — not dropped. A refusal is data.
Published with each release Report page · methodology page · raw CSV and JSON · the query set · the collection code · a changelog · a permanent version identifier.

One methodological choice deserves its own explanation: why 72 hours, specifically, rather than a single instant or a full month. A single instant risks catching a temporary blip — a model rollback, a brief outage, an unrelated product incident — and reporting it as the quarter's baseline. A full month reintroduces the drift problem the window is meant to control for, since a product can change meaningfully within thirty days. Seventy-two hours is a compromise: long enough to average out a single anomalous hour, short enough that "this is what the surface looked like in this window" stays a meaningful, falsifiable statement.

The variance problem

This deserves its own section, because it undermines a great deal of what is currently published as fact.

Language models are not deterministic. Ask the same question twice and you may get a different answer built from different sources. Almost every AI-visibility figure in circulation — including the ones charted above — comes from a single run per query. If the re-run rate of source disagreement is meaningful, then some proportion of every published citation statistic is measurement error rather than signal, and nobody currently knows what that proportion is.

Why five runs matters more than it sounds

Five runs per query does not eliminate the problem. It makes the problem visible. Every figure the Index publishes will carry the spread across runs alongside the central estimate, so a reader can see whether a difference between two engines is real or within noise. Quantifying that spread is hypothesis H2, and it is the first thing v1 will report.

Almost every AI-visibility number in circulation comes from a single query run. If AI answers aren't deterministic, some share of every published statistic is measurement error — and nobody knows how much.

Share on X

Consider a concrete, hypothetical illustration of why this matters practically. Suppose Domain A shows a 40% citation rate on a single-run measurement and Domain B shows 34%. A report built on that single run would confidently declare Domain A "more visible." But if five runs on Domain A actually range from 28% to 52%, and five runs on Domain B range from 30% to 38%, the honest conclusion is that the two domains are statistically indistinguishable — the single-run comparison manufactured a difference that the repeated measurement dissolves. This is not a hypothetical failure mode unique to this example; it is the default failure mode of every single-run citation study published to date, including several referenced elsewhere on this site.

Four metrics worth defining

New terminology should earn its place. Each of these exists because there is a real measurement problem with no current answer — not to coin a phrase.

CR

Citation Rate — the share of tracked queries in which a domain appears at least once. The base unit. Right now every vendor means something different by 'AI visibility'.

Defined here
SOV

AI Share of Voice — a domain's citations as a share of all citations in a query set. Makes competitive comparison possible on a fixed denominator.

Defined here
CHL

Citation Half-Life — days until a cited URL's citation rate falls to half its peak. Nobody has measured whether a citation persists at all.

Novel
CE

Citation Efficiency — citations earned per 1,000 pages crawled. Connects server-log reality to visibility outcome. No equivalent metric exists.

Novel

There is one metric deliberately not on that list: a composite "AI Visibility Score". Every visibility vendor has one, they are unfalsifiable, and building one here would undercut the only thing this project has going for it.

How the metrics are calculated, with a worked example

Definitions are easier to trust when the arithmetic behind them is visible. Below is each formula, followed by a single worked example using invented numbers for a hypothetical domain, "acmegear.com."

MetricFormula
Citation Rate (CR)queries citing the domain ÷ total tracked queries
Share of Voice (SOV)domain's citations ÷ all citations across the query set
Citation Half-Life (CHL)days elapsed when citation rate for a cohort of URLs first drops to 50% of its peak value
Citation Efficiency (CE)(citations earned ÷ pages crawled by that engine's bot) × 1,000
34%

acmegear.com: cited in 340 of the 1,000 tracked queries — a Citation Rate of 34%.

Worked example, invented figures

Continuing the same hypothetical: if acmegear.com earned 340 citations out of 9,800 total citations recorded across the whole query set that quarter, its Share of Voice would be 340 ÷ 9,800, or roughly 3.5% — a meaningfully different, and arguably more useful, number than the bare citation rate, because it accounts for how crowded the overall citation landscape was that quarter. If server logs show an AI bot crawled 12,000 of acmegear.com's pages that same quarter, its Citation Efficiency would be (340 ÷ 12,000) × 1,000 ≈ 28.3 citations per 1,000 pages crawled — a number that becomes genuinely interesting only once compared against a competitor's efficiency figure, since a site that earns the same citation count from a tenth of the crawl volume is doing something structurally different with its content.

Pre-registered hypotheses

These are posted now, before collection, with the direction predicted and the analysis specified. If the data contradicts them, that gets published as the result.

Hypotheses registered for v1 — September 2026
#HypothesisPredictionCurrent status
H1Cross-engine citation overlap is below 25% for the same querySupportedCorrelational precedent only
H2Repeated identical queries return different source sets in over 30% of casesSupportedUntested in public
H3Google top-10 rank predicts citation for AI Overviews but not for ChatGPT or ClaudeSupportedPartial precedent (Ahrefs)
H4Citation rate is more concentrated than organic ranking — fewer domains take a larger shareSupportedUnmeasured
H5A majority of cited URLs remain cited 90 days laterNot supportedNever measured
H6Domain rating correlates with citation rate at r > 0.4Not supportedWeak precedent suggests lower

H5 and H6 are the interesting ones, because the prediction is that they fail. Registering a hypothesis you expect to reject is the cheapest available protection against reading a pattern into noise after the fact — and if H6 does hold, that is a genuinely surprising result worth more than a confirmation.

How this compares to existing visibility tools

The Index is not trying to replace a paid visibility dashboard for day-to-day monitoring — those tools are built for a different job (continuous tracking of your own brand across many prompts) and do it well. The comparison below is about what each category is structurally able to offer, not about which product is "better."

PropertyCommercial visibility toolsThe AI Citation Index
Raw data downloadableNo — dashboard onlyYes, every release
Methodology fully disclosedPartial, varies by vendorPublished before collection
Continuous monitoring of your own brandYes — this is their core productNo — fixed public query set only
Cost to accessSubscriptionFree, CC BY 4.0
Multi-run variance reportedRarely disclosedAlways, per release
Reproducible by a third partyNoYes — code and query set are public

In practice, the two are complementary: a brand that wants to track its own daily citation rate across a custom prompt set still needs a commercial tool. A researcher, journalist, or another tool builder who needs a checkable, reproducible baseline for the field as a whole has, as of this writing, nowhere else to go — which is the specific need this project is built to serve.

What the Index will not prove

This matters as much as the design. The Index is an observational instrument, and observational data has hard limits that are easy to forget once a chart looks convincing.

  • It cannot establish causation. If highly-cited domains share a trait, that does not mean the trait produced the citations. Causal claims need randomised intervention, which is a separate programme.
  • It measures the engines, not your site. Citation rates on a fixed public query set describe the ecosystem. They do not tell an individual site owner what will happen to them.
  • It is a snapshot of moving targets. These products change without notice. A finding is true of the surface as it behaved during a stated 72-hour window, and every figure is published with that window attached.
  • It is English-first at v1. Multilingual collection is planned but not funded yet, and results should not be generalised beyond the language tested.
  • Sampling bias is real. A 1,000-query set is a choice, and a different set would produce different numbers. That is exactly why the set is published rather than described.
  • It cannot detect a citation an engine never surfaces to the collection environment specifically. If an engine personalises results even in a logged-out state based on IP geography or device fingerprint in ways the collection setup does not fully neutralise, the recorded citation set may not match what every real user sees.

How to check our work

An open study that nobody audits is just a closed study with extra steps. Three things are published specifically so the findings can be attacked:

  1. The query set. If the queries are badly chosen, the results are wrong, and you can see the queries. Disagreement about the set is legitimate criticism and gets published alongside it.
  2. The collection code. Extraction logic is where quiet errors live — a selector that misses a citation type will skew everything downstream. The code ships with each release.
  3. The raw records. Not summary tables. Every individual citation row, so anyone can re-run the analysis and reach a different conclusion from the same data.

If you find an error, it gets corrected in a dated changelog entry rather than silently edited — and the person who found it gets credited.

Release schedule

Quarterly full releases, monthly pulse updates in between. Each release is versioned and permanent — older versions stay online rather than being overwritten, so historical figures remain checkable.

Planned releases
  1. v0Sep 2026

    Protocol only — this pre-registration.

    Method committed before any collection.

  2. v2Q2 2027

    Expands to 5,000 queries.

    First per-vertical breakdowns.

  3. v3Q3 2027

    5,000 queries, third collection window.

    First citation half-life estimates become possible.

  4. v4Q4 2027

    Full-year synthesis.

    Twelve months of trend data and the annual report.

If a release slips, the delay gets posted here with the reason. If the project stops, that gets posted too. An abandoned research programme that quietly goes stale is worse than one that never started.

Contribute a query, or your logs

Two things would materially improve this work. If you have queries in your vertical that you think belong in the set, send them — the set is public and contributions are credited. And if you run a site and would share anonymised server logs, that unlocks the crawl-side measurements that the query data alone cannot reach. Both routes are open at hello@ritiknamdev.com.

Quick glossary

CitationAn AI-generated answer referencing or linking to a specific source URL.
Citation Rate (CR)Share of tracked queries where a domain is cited at least once.
Share of Voice (SOV)A domain's citations as a share of all citations in the query set.
Citation Half-Life (CHL)Days until a URL's citation rate falls to half its peak.
Citation Efficiency (CE)Citations earned per 1,000 pages an AI bot crawled from the site.
Collection windowThe fixed period (72 hours per quarter) during which a release's data is gathered.

For every other term used across this site's research — GEO, AEO, query fan-out, crawl-to-referral ratio and more — see the full AI search glossary.

How to cite this
Namdev, R. (2026). The AI Citation Index (pre-registration) (v0). Retrieved from https://ritiknamdev.com/blog/ai-citation-index

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

The measurement approach here grew out of two first-party studies: a 90-day server-log test of llms.txt, which is where the crawl-side metrics come from, and Zero to Cited, which tracked a brand-new domain into AI search. For the technical groundwork, see GPTBot vs OAI-SearchBot and the 50-point technical GEO audit. For where the numbers on this page came from and how confident to be in each, see where AI SEO statistics actually come from.

FAQ

Frequently asked questions

Why publish the methodology before collecting any data?
Because a hypothesis posted after the results are in is not a hypothesis — it is a description. Pre-registration is standard practice in science and almost entirely absent from SEO research. Posting the protocol first means a null result is as publishable as a positive one, which is the only way to keep the findings honest.
How is this different from Profound, Peec or the Semrush AI Toolkit?
Those are visibility products with proprietary corpora, and several of them are excellent. The difference is structural, not qualitative: their data is the product, so they cannot open-source it. The Index publishes the query set, the raw citation records and the collection code. You can re-run it and check the numbers.
Is a 1,000-query sample big enough?
For domain-level citation rates across seven engines, yes — the confidence intervals are reported alongside every figure. For long-tail or per-vertical questions it is not, which is why v2 expands to 5,000 queries. Any figure whose interval is too wide to support a claim will be published with that stated, not quietly dropped.
Why five runs per query instead of one?
Because AI answers are not deterministic. The same question can return different sources on a second attempt, so a single run measures noise as well as signal. Five runs allows the variance to be reported rather than hidden — and quantifying that variance is itself one of the first findings the Index will publish.
Can I use the data commercially?
Yes. Everything is CC BY 4.0 — reuse it, chart it, build on it, sell analysis of it. The only requirement is attribution back to the Index release you used, by version number.
What happens if the results contradict something you have already published?
The finding gets published anyway, along with a note about what changed and why. That includes contradicting my own earlier guides on this site.
Will the Index track paid placements or ads inside AI answers?
Not at v1. Sponsored placements inside AI answers are a newer, less standardised phenomenon across engines, and mixing them into an organic-citation dataset would muddy the metric before it even has a baseline. A separate tracker for sponsored placements is a candidate for a later version, tracked openly rather than folded in silently.
How do you handle an engine that changes its product mid-quarter?
The 72-hour collection window limits exposure to drift within a single release, but a major mid-quarter change (a new model version, a redesigned answer UI) gets logged in that release's changelog as a confound, not smoothed over. If the change is large enough to make the quarter's numbers unrepresentative, that gets stated plainly in the release notes rather than buried in a footnote.
Why not just use an existing academic benchmark instead of building a new query set?
Academic IR benchmarks are built for retrieval accuracy, not for citation-selection behaviour in commercial AI-search products, and none we reviewed cover all seven surfaces tracked here with a stratification aimed at SEO and GEO use cases specifically. Reusing the Princeton GEO paper's benchmark was considered, but at 10,000 queries built for a 2023-24 landscape, it predates AI Mode and Claude's web search tool entirely.
What stops this project from just becoming another vendor with a proprietary score?
Nothing except a standing commitment, which is why it's stated this explicitly: no composite score, raw data published every release, methodology first, and null results treated as findings rather than failures. If a future version of this page ever drops the raw-data commitment, that is the moment to stop trusting it — and the changelog will make any such change impossible to hide.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.