A vendor-reported figure — no public sample, no date, graded Partial here — puts ChatGPT/Perplexity domain overlap at roughly 11%: the two engines agree on a cited source about one time in nine. Independent work on AI search overlap points the same way without settling the number. This page registers a study to independently verify that figure, and extend it across all seven engines tracked by the Citation Index, broken down by query type. No queries have been run yet.
No data has been collected yet. Nothing on this page is a result. Design covers seven engines. No queries have been run under this protocol.
- A vendor-reported analysis puts ChatGPT/Perplexity domain overlap at roughly 11%. It is graded Partial here — no public sample or date.
- No rigorous cross-engine concordance measurement has been published by anyone.
- Actual agreement across all seven engines, rather than an estimate between two.
- Whether concordance is higher for narrow factual queries than broad subjective ones.
- Whether Google's two surfaces agree with each other more than either agrees with a rival.
Why does cross-engine agreement matter?
If AI search engines mostly agreed on what to cite for a given question, "AI visibility" could reasonably be treated as one single target. The available evidence points the other way: the qualitative accounts published so far describe noticeably different selection behaviour per engine, which would mean a strategy tuned to one engine's citation pattern transfers poorly to another's.
Quantifying how much disagreement exists — and whether it varies by query type — turns a widely repeated impression into something measurable. What GEO is covers the discipline around it.
What is actually known today?
- Cited by both ChatGPT and Perplexity 11
- Cited by only one 89
One figure, and it is Partial : vendor-reported, from a corpus that is not independently reproducible, with no public sample and no stated collection date. That is the same weakness the provenance audit finds in most numbers in this field, and the reason the measurement standard asks for the raw ingredients instead. It remains the single best available estimate today — which is precisely why a first-party, reproducible version is worth building rather than quoting this one for another year.
How is concordance calculated?
Cross-Platform Concordance is a Jaccard index: intersection over union, of cited domain sets between two engines, for the same query, averaged across the full query set. A score of 1.0 would mean two engines always cite an identical domain set; 0 would mean no overlap, ever.
Worked through on an invented, illustrative single query: for "best project management software for small teams," suppose ChatGPT cites four domains — A, B, C, D — and Perplexity cites five — B, C, E, F, G. The intersection is {B, C}, a count of 2. The union is {A, B, C, D, E, F, G}, a count of 7. The Jaccard index for that query is 2 ÷ 7, roughly 0.29. Averaged across the full 1,000-query set, that produces the concordance figure reported for each engine pair.
How the study is designed, and what it predicts
Concordance falls directly out of the AI Citation Index's core collection. Every engine already runs against the same published query set, so this needs no separate data collection — only a different analysis of data already being gathered, alongside the other protocols on the studies index.
| # | Hypothesis | Predicted outcome |
|---|---|---|
| CC1 | Overall concordance across all seven engines is below 20% | Expect to hold |
| CC2 | Concordance is higher for narrow, factual queries than for broad, subjective ones | Expect to hold |
| CC3 | Google's two surfaces (AI Overviews and AI Mode) show higher mutual concordance than either does with a non-Google engine | Expect to hold |
With seven engines, the eventual output is not a single number but a 7×7 matrix of pairwise scores — one cell per engine pair, readable at a glance for which pairs behave most alike. A matrix also tests CC3 directly, by comparing the AI-Overviews/AI-Mode cell against every cell involving a non-Google engine. A single blended average would hide exactly that kind of structure.
If AI search engines mostly agreed on what to cite, 'AI visibility' could be treated as one target. The best current estimate — vendor-reported, no public sample or date, graded Partial — is ~11% agreement between just two engines.
Share on XWhy would two engines disagree?
Each possible cause implies a different reading of the eventual number, so it is worth naming them before measuring anything.
Different corpora. Engines crawl different subsets of the web, at different depths, on different schedules. Two systems cannot agree about a page one of them never fetched, which makes crawler behaviour an upstream cause of any concordance figure. Seer's finding that 87% of SearchGPT citations matched Bing's top results is the clearest published example of a shared corpus producing shared answers.
Different licensing. Some engines have content agreements others do not. A licensed source can be over-represented for commercial reasons unrelated to quality.
Different retrieval methods. Keyword-style matching, embedding similarity and hybrid approaches surface different pages from the same corpus.
Different fan-out. If one engine decomposes a question into five sub-queries and another into two, they are effectively answering different questions. Google's granted patent describes eight sub-query categories, and later public comments confirm the technique is live in AI Mode. No equivalent documentation exists for Anthropic's or Perplexity's retrieval.
Different answer length. An engine that writes six paragraphs needs more sources than one that writes two. Longer answers mechanically raise the chance of overlap — a measurement artefact rather than a finding, and the reason any concordance figure has to be read alongside how many sources each engine used, or it silently measures verbosity.
Domain, URL or claim: which agreement?
"Do these engines agree?" is three questions wearing one coat, and they give three different numbers. We intend to report the first two numerically and treat the third as a qualitative annotation on a subsample; pretending claim-level agreement can be automated would overstate what the method can do.
| Level | Question it answers | Expected reading | How we report it |
|---|---|---|---|
| Domain | Do both engines cite the same publishers? | Highest of the three — the most forgiving measure | Numerically, full panel |
| URL | Do both cite the same specific pages? | Lower than domain level, usually substantially | Numerically, full panel |
| Claim | Do the answers actually assert the same things? | Unknown — two engines can cite nothing in common and still agree | Qualitative annotation on a subsample only |
Why query type should dominate the result
Our expectation, stated before collection, is that query type explains more of the variance than engine identity does. Stable factual questions have a small set of obvious authoritative sources, so every engine should converge and concordance should be high but uninformative. Commercial comparison questions have no canonical source, a large candidate pool and strong incentives for many parties to publish, so concordance should be low. Current-events questions depend on crawl freshness, which differs sharply between engines, so concordance should be low and unstable across runs. How-to questions sit somewhere between.
If that holds, a single site-wide concordance number is close to meaningless and every published figure should be broken out by query type. That would itself be a useful result, and one we would report even though it makes our headline number less quotable.
How much does an engine agree with itself?
Every concordance study needs a control that almost none of them include. Generated answers vary between runs: ask the same engine the same question twice and the source list may differ. If self-agreement is, say, seventy percent, then cross-engine agreement of eleven percent — the vendor-reported figure, no public sample or date, graded Partial — means something very different than it would if self-agreement were ninety-eight percent.
Without that control, a concordance figure conflates genuine differences between engines with ordinary run-to-run randomness. The first is interesting; the second is noise reported as a finding. We intend to measure the noise floor first, publish it alongside every cross-engine number, and register the outcome in the null results registry whichever way it falls. If the floor turns out to be low enough that engines barely agree with themselves, the honest conclusion is that the metric does not work, and we will say so.
How should you read any concordance number?
Four questions decide whether a cross-platform agreement statistic means anything — ours included. A figure that cannot answer all four is a talking point, not a measurement, which is why this design is published before the data. The method page describes that discipline in full.
One borrowed idea is worth stating plainly: inter-rater reliability in the social sciences corrects for agreement that would happen by chance. Two engines both citing the single obvious source for a factual question is not evidence of similar retrieval. A raw overlap percentage is the beginning of the analysis, not the end of it.
What would a high or low result mean for you?
Underneath the methodology there is one commercial question: if I invest in being visible in one AI engine, do I get the others for free? Nobody can currently answer that with evidence, which has not stopped it being answered confidently in both directions. Vendors selling per-platform tooling have an interest in the answer being no; agencies selling one retainer have an interest in it being yes.
If concordance is high, effort transfers between engines, per-platform optimisation is mostly wasted, and the sensible strategy is one body of well-sourced work published once. That is roughly the position the one controlled GEO study supports, since its mechanisms are not engine-specific.
If concordance is low, being visible in one engine tells you almost nothing about the others, measurement has to be per-engine, and a single "AI visibility" score is actively misleading. It still would not follow that you need separate content — only separate measurement, a distinction most vendors have a commercial reason to blur. A page cited heavily by Perplexity but ignored by ChatGPT is a genuinely different result from a page ignored by both, and one blended total cannot tell those apart.
If concordance varies mostly by query type, which is our expectation, the answer is neither: you would allocate by where your questions sit, not by engine.
Open question Until the study runs, the honest position for a practitioner is to assume partial transfer, invest in fundamentals that plausibly help everywhere, and measure each engine separately rather than assuming either extreme. If the final numbers do land near the vendor-reported ~11% — no public sample or date, graded Partial — the first practical consequence is to stop optimizing against a single composite "AI visibility" score.
Why repeat it instead of running it once?
A single concordance figure has a short shelf life. These products change often, and without telling anyone. That makes a one-off study almost useless a year later — and worse, it keeps getting quoted, because a number with no expiry date circulates forever.
The useful version is a repeated measurement on a fixed panel: the panel stays the same, the date changes, and the comparison between runs is where the information lives. That is why concordance is designed as part of the Citation Index rather than as a standalone piece. It also creates an obligation — once you publish a tracked figure you have to keep publishing it, including in the quarters where nothing interesting happened.
How could this study fail, and what would we publish then?
Listing the failure modes in advance is part of the design. Each has a mitigation, and each mitigation costs something.
Personalisation contaminating results. Engines may tailor answers to account history or location. Mitigation: clean sessions, fixed region, conditions recorded on every run. Cost: the results describe a fresh anonymous user, not a typical one.
Interface drift mid-study. A product can change while collection is running. Mitigation: collect all engines within a short window and note any observed change. Cost: a narrower window means a smaller panel.
Panel bias. Our own topical interests would skew a query set toward marketing questions. Mitigation: build the panel from public question sources across several categories, and publish it. Cost: some queries will be outside our expertise to interpret.
Over-reading a small panel. The most likely failure, and the most tempting. Mitigation: publish confidence intervals, and refuse to report a headline number the sample cannot support.
The null results are stated in advance too, because a study that can only produce an interesting result is not a study. If concordance is indistinguishable from the noise floor, we publish that the metric fails. If query type does not explain the variance, that goes on the page next to the prediction it falsified. If the panel is too small to separate the engines, we report the confidence interval and decline to publish a headline number at all. And if the result simply reproduces what is already published, we say so plainly rather than dressing a replication as a discovery.
Why these seven engines?
The point of seven is not completeness but coverage of distinct architectures and audiences. Two engines that behave identically add cost and no information.
| Engine | Why it is in the panel | What it makes measurable |
|---|---|---|
| ChatGPT Search | Largest assistant by usage | Runs its own retrieval crawler, so access is observable in server logs |
| Perplexity | Presents itself most visibly as a search product | Lists sources explicitly — the cheapest engine to measure accurately |
| Google AI Overviews | Widest passive reach; most people who see it did not choose an AI product | Citation behaviour on users who never opted into an assistant |
| Google AI Mode | A separate surface with separate behaviour | Whether two surfaces from one company agree — folding it into AI Overviews is the most common error in this field |
| Gemini | The standalone assistant, distinct again | Behaviour nobody has measured against the two Google search surfaces |
| Microsoft Copilot | Sits on a different index and is embedded in workplace software | Whether a differently-skewed user base changes citation share at all |
| Claude | Usage skews technical and professional | Whether audience shapes retrieval behaviour — this is where a difference should show up |
Newer surfaces — agentic browsers, MCP servers, WebMCP — are excluded because they do not yet produce a comparable citation list.
Running a small version yourself
You do not need our study to answer the version of this question that affects you. A twenty-query panel across three engines is an afternoon of work. Before any of it, check that each engine can reach you at all: the technical GEO audit and the JavaScript rendering question cover the access layer.
- 1 Use your own questions The ones your buyers actually ask. A generic panel answers a generic question you do not have.
- 2 Record every source Full source list per engine per query. Domains at minimum, URLs if you can manage it.
- 3 Run each query twice On the same engine, for your own noise floor. The step everyone skips and the one that makes the rest interpretable.
- 4 Compute intersection over union Per query, then look at the distribution rather than the average. A few queries usually behave completely differently.
- 5 Date it and repeat One number tells you where you stand. Two tell you which way things are moving.
Limitations
- Concordance measures agreement, not quality — two engines could agree on a poor source, or disagree while both citing good ones.
- The result depends heavily on the query set — a different set weighted toward more subjective or more factual queries could shift the overall figure meaningfully.
- The only prior figure is vendor-reported — the ~11% ChatGPT/Perplexity overlap has no public sample and no stated date, and is graded Partial throughout this page. It is a starting point to verify, not a baseline to trust.
Results and the raw per-query files go out to the newsletter when the first collection window closes. If you want to propose a query for the panel or point out a flaw in the design before it runs, the contact address is on the about page.
Namdev, R. (2026). Cross-platform citation concordance (v1). Retrieved from https://ritiknamdev.com/blog/cross-platform-citation-concordance Published under CC BY 4.0 — reuse freely with attribution.
Derived directly from the AI Citation Index's shared query set — see most-cited domains in AI search for the existing (vendor-reported) two-engine overlap figure this study aims to independently verify and extend.