Original research · Pre-registered

Cross-platform citation concordance

How much do ChatGPT, Perplexity, Gemini, Claude, and Google's AI surfaces actually agree on what to cite for the same question? A vendor-reported ~11% overlap between two engines — no public sample or date, graded Partial — suggests the answer is: not much.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — design stage ·13 min read ·Last verified September 2026
The short version

A vendor-reported figure — no public sample, no date, graded Partial here — puts ChatGPT/Perplexity domain overlap at roughly 11%: the two engines agree on a cited source about one time in nine. Independent work on AI search overlap points the same way without settling the number. This page registers a study to independently verify that figure, and extend it across all seven engines tracked by the Citation Index, broken down by query type. No queries have been run yet.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. Design covers seven engines. No queries have been run under this protocol.

What is known
  • A vendor-reported analysis puts ChatGPT/Perplexity domain overlap at roughly 11%. It is graded Partial here — no public sample or date.
  • No rigorous cross-engine concordance measurement has been published by anyone.
What is not yet known
  • Actual agreement across all seven engines, rather than an estimate between two.
  • Whether concordance is higher for narrow factual queries than broad subjective ones.
  • Whether Google's two surfaces agree with each other more than either agrees with a rival.

Why does cross-engine agreement matter?

If AI search engines mostly agreed on what to cite for a given question, "AI visibility" could reasonably be treated as one single target. The available evidence points the other way: the qualitative accounts published so far describe noticeably different selection behaviour per engine, which would mean a strategy tuned to one engine's citation pattern transfers poorly to another's.

Quantifying how much disagreement exists — and whether it varies by query type — turns a widely repeated impression into something measurable. What GEO is covers the discipline around it.

What is actually known today?

Reported domain overlap, ChatGPT vs. Perplexity citations
  • Cited by both ChatGPT and Perplexity 11
  • Cited by only one 89
Source: Profound, vendor-reported from a non-public corpus — graded partial. Only two engines compared; the full seven-engine picture is what this study will measure independently.

One figure, and it is Partial : vendor-reported, from a corpus that is not independently reproducible, with no public sample and no stated collection date. That is the same weakness the provenance audit finds in most numbers in this field, and the reason the measurement standard asks for the raw ingredients instead. It remains the single best available estimate today — which is precisely why a first-party, reproducible version is worth building rather than quoting this one for another year.

How is concordance calculated?

Cross-Platform Concordance is a Jaccard index: intersection over union, of cited domain sets between two engines, for the same query, averaged across the full query set. A score of 1.0 would mean two engines always cite an identical domain set; 0 would mean no overlap, ever.

Worked through on an invented, illustrative single query: for "best project management software for small teams," suppose ChatGPT cites four domains — A, B, C, D — and Perplexity cites five — B, C, E, F, G. The intersection is {B, C}, a count of 2. The union is {A, B, C, D, E, F, G}, a count of 7. The Jaccard index for that query is 2 ÷ 7, roughly 0.29. Averaged across the full 1,000-query set, that produces the concordance figure reported for each engine pair.

How the study is designed, and what it predicts

Concordance falls directly out of the AI Citation Index's core collection. Every engine already runs against the same published query set, so this needs no separate data collection — only a different analysis of data already being gathered, alongside the other protocols on the studies index.

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
CC1Overall concordance across all seven engines is below 20%Expect to hold
CC2Concordance is higher for narrow, factual queries than for broad, subjective onesExpect to hold
CC3Google's two surfaces (AI Overviews and AI Mode) show higher mutual concordance than either does with a non-Google engineExpect to hold

With seven engines, the eventual output is not a single number but a 7×7 matrix of pairwise scores — one cell per engine pair, readable at a glance for which pairs behave most alike. A matrix also tests CC3 directly, by comparing the AI-Overviews/AI-Mode cell against every cell involving a non-Google engine. A single blended average would hide exactly that kind of structure.

If AI search engines mostly agreed on what to cite, 'AI visibility' could be treated as one target. The best current estimate — vendor-reported, no public sample or date, graded Partial — is ~11% agreement between just two engines.

Share on X

Why would two engines disagree?

Each possible cause implies a different reading of the eventual number, so it is worth naming them before measuring anything.

Different corpora. Engines crawl different subsets of the web, at different depths, on different schedules. Two systems cannot agree about a page one of them never fetched, which makes crawler behaviour an upstream cause of any concordance figure. Seer's finding that 87% of SearchGPT citations matched Bing's top results is the clearest published example of a shared corpus producing shared answers.

Different licensing. Some engines have content agreements others do not. A licensed source can be over-represented for commercial reasons unrelated to quality.

Different retrieval methods. Keyword-style matching, embedding similarity and hybrid approaches surface different pages from the same corpus.

Different fan-out. If one engine decomposes a question into five sub-queries and another into two, they are effectively answering different questions. Google's granted patent describes eight sub-query categories, and later public comments confirm the technique is live in AI Mode. No equivalent documentation exists for Anthropic's or Perplexity's retrieval.

Different answer length. An engine that writes six paragraphs needs more sources than one that writes two. Longer answers mechanically raise the chance of overlap — a measurement artefact rather than a finding, and the reason any concordance figure has to be read alongside how many sources each engine used, or it silently measures verbosity.

Domain, URL or claim: which agreement?

"Do these engines agree?" is three questions wearing one coat, and they give three different numbers. We intend to report the first two numerically and treat the third as a qualitative annotation on a subsample; pretending claim-level agreement can be automated would overstate what the method can do.

LevelQuestion it answersExpected readingHow we report it
DomainDo both engines cite the same publishers?Highest of the three — the most forgiving measureNumerically, full panel
URLDo both cite the same specific pages?Lower than domain level, usually substantiallyNumerically, full panel
ClaimDo the answers actually assert the same things?Unknown — two engines can cite nothing in common and still agreeQualitative annotation on a subsample only

Why query type should dominate the result

Our expectation, stated before collection, is that query type explains more of the variance than engine identity does. Stable factual questions have a small set of obvious authoritative sources, so every engine should converge and concordance should be high but uninformative. Commercial comparison questions have no canonical source, a large candidate pool and strong incentives for many parties to publish, so concordance should be low. Current-events questions depend on crawl freshness, which differs sharply between engines, so concordance should be low and unstable across runs. How-to questions sit somewhere between.

If that holds, a single site-wide concordance number is close to meaningless and every published figure should be broken out by query type. That would itself be a useful result, and one we would report even though it makes our headline number less quotable.

How much does an engine agree with itself?

Every concordance study needs a control that almost none of them include. Generated answers vary between runs: ask the same engine the same question twice and the source list may differ. If self-agreement is, say, seventy percent, then cross-engine agreement of eleven percent — the vendor-reported figure, no public sample or date, graded Partial — means something very different than it would if self-agreement were ninety-eight percent.

Without that control, a concordance figure conflates genuine differences between engines with ordinary run-to-run randomness. The first is interesting; the second is noise reported as a finding. We intend to measure the noise floor first, publish it alongside every cross-engine number, and register the outcome in the null results registry whichever way it falls. If the floor turns out to be low enough that engines barely agree with themselves, the honest conclusion is that the metric does not work, and we will say so.

How should you read any concordance number?

Four questions decide whether a cross-platform agreement statistic means anything — ours included. A figure that cannot answer all four is a talking point, not a measurement, which is why this design is published before the data. The method page describes that discipline in full.

What was the unit?Domain, URL or claim. Quoting a domain-level figure as URL-level roughly doubles the apparent agreement.
What was the panel?Twenty consumer queries and five hundred mixed-intent queries are not comparable results.
Was the noise floor measured?Without a same-engine control you cannot tell disagreement from ordinary run-to-run randomness.
How many sources per answer?Longer answers mechanically overlap more. Without this, the figure silently measures verbosity.

One borrowed idea is worth stating plainly: inter-rater reliability in the social sciences corrects for agreement that would happen by chance. Two engines both citing the single obvious source for a factual question is not evidence of similar retrieval. A raw overlap percentage is the beginning of the analysis, not the end of it.

What would a high or low result mean for you?

Underneath the methodology there is one commercial question: if I invest in being visible in one AI engine, do I get the others for free? Nobody can currently answer that with evidence, which has not stopped it being answered confidently in both directions. Vendors selling per-platform tooling have an interest in the answer being no; agencies selling one retainer have an interest in it being yes.

If concordance is high, effort transfers between engines, per-platform optimisation is mostly wasted, and the sensible strategy is one body of well-sourced work published once. That is roughly the position the one controlled GEO study supports, since its mechanisms are not engine-specific.

If concordance is low, being visible in one engine tells you almost nothing about the others, measurement has to be per-engine, and a single "AI visibility" score is actively misleading. It still would not follow that you need separate content — only separate measurement, a distinction most vendors have a commercial reason to blur. A page cited heavily by Perplexity but ignored by ChatGPT is a genuinely different result from a page ignored by both, and one blended total cannot tell those apart.

If concordance varies mostly by query type, which is our expectation, the answer is neither: you would allocate by where your questions sit, not by engine.

Open question Until the study runs, the honest position for a practitioner is to assume partial transfer, invest in fundamentals that plausibly help everywhere, and measure each engine separately rather than assuming either extreme. If the final numbers do land near the vendor-reported ~11% — no public sample or date, graded Partial — the first practical consequence is to stop optimizing against a single composite "AI visibility" score.

Why repeat it instead of running it once?

A single concordance figure has a short shelf life. These products change often, and without telling anyone. That makes a one-off study almost useless a year later — and worse, it keeps getting quoted, because a number with no expiry date circulates forever.

The useful version is a repeated measurement on a fixed panel: the panel stays the same, the date changes, and the comparison between runs is where the information lives. That is why concordance is designed as part of the Citation Index rather than as a standalone piece. It also creates an obligation — once you publish a tracked figure you have to keep publishing it, including in the quarters where nothing interesting happened.

How could this study fail, and what would we publish then?

Listing the failure modes in advance is part of the design. Each has a mitigation, and each mitigation costs something.

Personalisation contaminating results. Engines may tailor answers to account history or location. Mitigation: clean sessions, fixed region, conditions recorded on every run. Cost: the results describe a fresh anonymous user, not a typical one.

Interface drift mid-study. A product can change while collection is running. Mitigation: collect all engines within a short window and note any observed change. Cost: a narrower window means a smaller panel.

Panel bias. Our own topical interests would skew a query set toward marketing questions. Mitigation: build the panel from public question sources across several categories, and publish it. Cost: some queries will be outside our expertise to interpret.

Over-reading a small panel. The most likely failure, and the most tempting. Mitigation: publish confidence intervals, and refuse to report a headline number the sample cannot support.

The null results are stated in advance too, because a study that can only produce an interesting result is not a study. If concordance is indistinguishable from the noise floor, we publish that the metric fails. If query type does not explain the variance, that goes on the page next to the prediction it falsified. If the panel is too small to separate the engines, we report the confidence interval and decline to publish a headline number at all. And if the result simply reproduces what is already published, we say so plainly rather than dressing a replication as a discovery.

Why these seven engines?

The point of seven is not completeness but coverage of distinct architectures and audiences. Two engines that behave identically add cost and no information.

The seven engines in the panel and the reason each earns a column. Selection rationale, fixed before collection — not a measurement of any engine's behaviour, and not a ranking.
EngineWhy it is in the panelWhat it makes measurable
ChatGPT SearchLargest assistant by usageRuns its own retrieval crawler, so access is observable in server logs
PerplexityPresents itself most visibly as a search productLists sources explicitly — the cheapest engine to measure accurately
Google AI OverviewsWidest passive reach; most people who see it did not choose an AI productCitation behaviour on users who never opted into an assistant
Google AI ModeA separate surface with separate behaviourWhether two surfaces from one company agree — folding it into AI Overviews is the most common error in this field
GeminiThe standalone assistant, distinct againBehaviour nobody has measured against the two Google search surfaces
Microsoft CopilotSits on a different index and is embedded in workplace softwareWhether a differently-skewed user base changes citation share at all
ClaudeUsage skews technical and professionalWhether audience shapes retrieval behaviour — this is where a difference should show up

Newer surfaces — agentic browsers, MCP servers, WebMCP — are excluded because they do not yet produce a comparable citation list.

Running a small version yourself

You do not need our study to answer the version of this question that affects you. A twenty-query panel across three engines is an afternoon of work. Before any of it, check that each engine can reach you at all: the technical GEO audit and the JavaScript rendering question cover the access layer.

A twenty-query concordance check you can run in an afternoon
  1. 1 Use your own questions The ones your buyers actually ask. A generic panel answers a generic question you do not have.
  2. 2 Record every source Full source list per engine per query. Domains at minimum, URLs if you can manage it.
  3. 3 Run each query twice On the same engine, for your own noise floor. The step everyone skips and the one that makes the rest interpretable.
  4. 4 Compute intersection over union Per query, then look at the distribution rather than the average. A few queries usually behave completely differently.
  5. 5 Date it and repeat One number tells you where you stand. Two tell you which way things are moving.

Limitations

  • Concordance measures agreement, not quality — two engines could agree on a poor source, or disagree while both citing good ones.
  • The result depends heavily on the query set — a different set weighted toward more subjective or more factual queries could shift the overall figure meaningfully.
  • The only prior figure is vendor-reported — the ~11% ChatGPT/Perplexity overlap has no public sample and no stated date, and is graded Partial throughout this page. It is a starting point to verify, not a baseline to trust.
When this runs

Results and the raw per-query files go out to the newsletter when the first collection window closes. If you want to propose a query for the panel or point out a flaw in the design before it runs, the contact address is on the about page.

How to cite this
Namdev, R. (2026). Cross-platform citation concordance (v1). Retrieved from https://ritiknamdev.com/blog/cross-platform-citation-concordance

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Derived directly from the AI Citation Index's shared query set — see most-cited domains in AI search for the existing (vendor-reported) two-engine overlap figure this study aims to independently verify and extend.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Discovered Labs — AI citation patterns: how ChatGPT, Claude and Perplexity choose sourcesdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Discovered Labs — How each platform cites sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Leapd — How ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Ahrefs — How much do AI search engines overlap with each other?ahrefs.com/blog/ai-search-overlap Ahrefs — AI Overview citations and the top 10ahrefs.com/blog/ai-overview-citations-top-10 Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Semrush — AI Overviews studywww.semrush.com/blog/semrush-ai-overviews-study Search Engine Journal — AI Overview citations from top-ranking pages drop sharplywww.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637 Search Engine Journal — Query fan-out in AI Mode: new details from Googlewww.searchenginejournal.com/query-fan-out-technique-in-ai-mode-new-details-from-google/552532 Google Patents — US11663201B2, query fan-out and sub-query categoriespatents.google.com/patent/US11663201B2 Google Search Central — AI features and your websitedevelopers.google.com/search/docs/appearance/ai-features Ziptie — How does ChatGPT choose its sources?ziptie.dev/blog/how-does-chatgpt-choose-its-sources Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors Anthropic — Web search tool documentationplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Perplexity — Developer documentationdocs.perplexity.ai Margen — Perplexity statisticswww.margen.net/perplexity-statistics-2026 Search Engine Journal — Google vs Microsoft Bingwww.searchenginejournal.com/google-vs-microsoft-bing/400855 Backlinko — Bing users and usage statisticsbacklinko.com/bing-users Statista — Bing worldwide market sharewww.statista.com/statistics/1219326/market-share-held-by-bing-worldwide Similarweb — Generative AI usage statisticsaisearch.similarweb.com/blog/gen-ai-stats GEO: Generative Engine Optimization — Aggarwal et al., KDD 2024arxiv.org/abs/2311.09735
FAQ

Frequently asked questions

Is the ~11% overlap unusually low?
There is no established benchmark to compare it against yet, and the figure itself is vendor-reported with no public sample or date behind it, graded Partial here. Nobody has published a rigorous cross-platform concordance study, with a disclosed method, covering more than two engines. That missing benchmark is exactly what this study is designed to establish.
Does low concordance mean AI-search strategy has to be built per-platform?
It is the strongest available argument for that approach. If engines genuinely agree only about 11% of the time — a vendor-reported figure with no public sample or date, graded Partial — then one unified "AI SEO" strategy may work far worse than platform-specific ones. This study aims to quantify how true that is, across all seven tracked engines, and by query type.
Why use a Jaccard index specifically, rather than a simpler percentage-overlap measure?
A Jaccard index, intersection over union, handles it correctly when two engines cite different numbers of total domains for the same query. A simple percentage-of-one-engine's-citations measure would be asymmetric. It would give a different number depending on which engine you treated as the baseline. The Jaccard index avoids that problem.
Why publish the study design before running the study?
Because a design published afterwards can be quietly reshaped to fit whatever the data turned out to say. Pre-registering the metric, the panel and the predictions is the cheapest available protection against fooling ourselves, and it lets anyone check that we did not move the goalposts.
What would make you abandon this study?
If the noise floor turns out to be as large as the between-engine differences. If the same engine, asked the same question twice, agrees with itself no more than it agrees with a competitor, then concordance is not measuring anything about engines and we would publish that finding and stop.
Can I use concordance to decide which engine to focus on?
Indirectly. Concordance tells you whether effort transfers between engines. It does not tell you which engine your buyers use. Pair it with your own audience evidence; on its own it answers the wrong half of the question.
Is domain-level or URL-level concordance the right measure?
They answer different questions and we intend to report both. Domain-level tells you whether the same publishers win everywhere. URL-level tells you whether the same specific pages do. The second is almost always lower, and confusing the two makes agreement look higher than it is.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.