Each major engine has a distinct centre of gravity: ChatGPT toward Wikipedia and encyclopedic sources, Perplexity toward Reddit, Google AI Overviews toward YouTube and multimodal content. Roughly 11% of domains are reportedly cited by both ChatGPT and Perplexity. All of these figures come from one vendor-reported corpus of a stated 680 million citations whose query mix is not public, so they are graded Partial throughout. Read the table per engine; a pooled "most-cited domains" list is the one thing this data cannot support.
- 47.9% of ChatGPT top citations reportedly go to Wikipedia and similarly encyclopedic sources — measured on a vendor corpus of a stated 680M citations, query mix undisclosed. Graded partial.
- 46.7% of Perplexity top citations reportedly go to Reddit, from the same corpus and the same undisclosed query mix. The two figures are comparable to each other and to nothing else.
- 23.3% of Google AI Overview citations reportedly favour YouTube and multimodal content — a lower concentration than either of the above, on the same corpus.
- Only ~11% of domains are reportedly cited by both ChatGPT and Perplexity. That overlap figure, not the source shares, is the finding that changes how work is planned.
- Every percentage on this page traces to one corpus nobody outside the vendor can inspect. The query set drives a citation-share table more than anything else does, and the query set is exactly what is not published.
Which domains does each engine actually cite?
of ChatGPT's top citations reportedly go to Wikipedia and similarly encyclopedic sources — the largest single-source concentration reported for any major engine.
of Perplexity's top citations reportedly go to Reddit, on the same corpus and the same undisclosed query mix.
of Google AI Overview citations reportedly favour YouTube and multimodal content — a notably lower concentration than the other two.
| Engine | Reported dominant source type | Population behind the figure | Structural implication for a publisher | Grade |
|---|---|---|---|---|
| ChatGPT | Wikipedia and encyclopedic reference — 47.9% of top citations | Vendor corpus, stated 680M citations; query mix, date range and counting rule undisclosed | Neutral, comprehensive, non-promotional coverage | Partial |
| Perplexity | Reddit and community discussion — 46.7% of top citations | Same corpus, same undisclosed query mix | Specific, first-hand, frequently updated detail | Partial |
| Google AI Overviews | YouTube and multimodal — 23.3% of citations | Same corpus; concentration roughly half that of the other two | A genuine visual or video companion, not a token one | Partial |
| ChatGPT (index lineage) | A large majority of SearchGPT citations matched Bing's top results | Seer Interactive's own sample; method described, data not published | Bing visibility is a load-bearing input most teams never check | Partial |
| Cross-engine (vs. classic Google) | About 12% of AI-cited URLs rank in Google's top 10 | Ahrefs, large URL sample; method described publicly | Classic rank is a weak predictor of citation | Traceable |
| Claude | No comparable per-engine share published | — | Treat any Claude source-mix claim as unmeasured | Broken chain |
Per-engine pages carry the detail: ChatGPT, Perplexity, Google AI Overviews, Claude and Gemini. What each dominant source stands in for has its own studies: the Wikipedia dependency study asks what ChatGPT cites when the encyclopedia has nothing on a topic, and the Reddit dependency study asks the same of Perplexity. Both are gap studies, and the gaps are where an ordinary publisher can actually win.
Why a single pooled most-cited list is the wrong asset
The genre convention is a ranked list of domains under a heading like "the 50 most-cited sites in AI search." That list cannot be built honestly from this data. Building it anyway is the weakness in almost every version you will find.
Pooling requires that each engine contributed a comparable sample. Nothing guarantees that here. The engines surface different numbers of sources per answer, so an engine that shows fifteen citations contributes more rows than one that shows three, from identical retrieval behaviour. The counting rule — per answer, per query, or per domain — is not stated, and each produces a different number from the same data. And a pooled share is a property of the query set as much as of the engines: change the questions and the table changes with no engine behaviour altered at all.
A pooled figure therefore describes a population that does not exist — a composite engine nobody uses. The per-engine rows above describe real, if unverifiable, populations. That is why the table is the asset on this page and the ranked list is not.
ChatGPT leans on Wikipedia. Perplexity leans on Reddit. Google AI Overviews lean on YouTube. Only ~11% of domains are cited by both ChatGPT and Perplexity. 'AI visibility' isn't one target — it's at least three.
Share on XHow much do the engines agree? About 11%
- Cited by both ChatGPT and Perplexity 11
- Cited by only one 89
An ~11% overlap means optimising for one engine's citation pattern gives you very little assurance of appearing in another's. This is the strongest available argument for measuring visibility per platform rather than as one blended score, consistent with the measurement standard.
Read it carefully, though. It says most domains cited by one engine are not cited by the other. It does not say the engines disagree about facts, or that one is better. It is not a probability for your specific site — a site with strong coverage of a topic may well be cited everywhere, and a weak one nowhere. Overlap also depends on how many domains each engine cites in total. If one draws from a much wider pool, low overlap is partly a consequence of pool size rather than of disagreement, and the pool sizes are not published either.
Overlap with classic Google rankings is a separate question, measured by Ahrefs across a large URL sample at about 12% of AI-cited URLs ranking in the top 10. Overlap between engines is the subject of the cross-platform citation concordance study, which is the work this section is really waiting on.
of ChatGPT's top citations reportedly go to Wikipedia and similarly encyclopedic sources — the largest single-source concentration reported for any major engine.
Why domain type, not domain authority, predicts citation
The pattern is best explained by retrieval architecture, not by any property of the cited domains. ChatGPT Search has historically leaned on a Bing-derived index with a bias toward broad, structured reference content. Seer Interactive found a large majority of SearchGPT citations matched Bing's top results, which makes the Microsoft index load-bearing for a product most people never associate with it. The Bing and Copilot guide covers that surface directly. Perplexity's own crawler and citation-first interface are associated with community-vetted, frequently updated discussion. Google's AI Overviews can draw on YouTube and the Knowledge Graph, inputs the other two do not have natively.
None of this is primarily about backlink profiles. The claim that backlink-derived authority matters less than brand presence is examined in brand mentions versus backlinks, with Zyppy's correlational work among the few public attempts to quantify any of it.
Two structural factors get asserted far more often than tested. First, structured data, where Ahrefs found no clear effect on its own corpus and our schema study takes it up; and named, credentialed authorship, the subject of the author E-E-A-T study. Every tactic in that family is graded in the tactic evidence scoreboard.
A skew can also arise for four different reasons that look identical from outside, and they imply different actions.
| Mechanism | What it means | Can a publisher act on it? |
|---|---|---|
| Index composition | The engine's underlying index simply contains more of that source type. | Only by being in that index at all. |
| Retrieval weighting | The ranking step favours certain structural or authority signals. | Partly — structure is controllable. |
| Availability | The source is openly licensed, unpaywalled and easy to fetch. | Yes — access is a choice, and the least discussed one. |
| Query mix | The measured questions happen to suit that source type. | Not at all — this is a measurement artefact. |
The availability row is the most actionable. A source that is blocked or hard to fetch cannot be cited whatever its quality, and that is a configuration decision rather than a content one. Each operator publishes its own agents — OpenAI, Perplexity and Anthropic — consolidated in the bot user-agent registry. Whether those crawlers render JavaScript decides whether your content exists to them at all, and the technical audit is the checklist version. Cloudflare's report on undeclared crawlers is a reminder that the declared list and the actual traffic are different objects.
Why every number here is graded partial
A partial grade means the figure traces to a named source that describes its own data, but the data itself cannot be inspected by anyone outside that source. We can name who said it. We cannot check it.
Four things specifically are missing, and each of them moves the number. The query set, which drives everything: a corpus weighted toward general-knowledge questions shows encyclopedic dominance almost by construction, and one weighted toward product research does not. The date range, since retrieval systems change without announcement. The country and language mix. And the definition of a "top citation" — counted once per answer, once per query, or once per domain, which changes the arithmetic substantially.
None of this makes the figures wrong. Unverifiable is a distinct category from wrong. The direction here is corroborated in independent write-ups rather than independent data: Discovered Labs on how each assistant chooses sources, Leapd's cross-platform summary and Ziptie on ChatGPT's source selection all describe the same lean without re-measuring it. Corroboration and replication are different things. We report the direction confidently and the decimals cautiously.
What bends a most-cited-domains result
If you are evaluating someone else's measurement, or running your own, these are the levers. Query set composition dominates everything else. Personalisation and location mean a corpus collected in one region describes that region. Time makes every figure a snapshot — how quickly a new page enters the citable pool is the crawl-to-citation latency question, and how long a citation survives is the half-life question.
Counting rules and deduplication of subdomains and regional variants both change concentration figures noticeably. And sampling of answers matters: engines can answer the same question differently on repeated attempts, and no published figure we have seen reports that variance.
Building your own citation sample
A small first-party sample beats a large corpus you cannot inspect, provided you are honest about its limits.
- 01 Freeze a question list Thirty to fifty real customer questions, in natural language
- 02 Fix the counting rule Decide per-answer, per-query or per-domain before you start
- 03 Run every engine On at least three separate days, because single samples hide variance
- 04 Record everything Including refusals and answers with no citations at all
- 05 Tabulate by type The domain-type table is more instructive than the domain table
- 06 Publish the method To yourself at minimum — otherwise it is a snapshot, not a measurement
Freeze thirty to fifty questions a real customer would ask, in natural language rather than keywords. Decide the counting rule before you start — we suggest recording every cited domain once per answer and keeping the raw answer text. Run each question on each engine you care about, across at least three separate days.
Record everything, including questions that produced no citations and answers that refused; dropping those biases the result. Tabulate by domain and by domain type, because the type table tells you what to build rather than who to envy. Then write down the method — the query list, the dates, the rule. If you cannot repeat the exercise in three months and get a comparable number, you built a snapshot rather than a measurement. The zero-to-cited log study is our own worked version of this.
Common misreadings of a citation-share table
"Half of citations go to one site, so the rest of us are locked out." A plurality is not a monopoly. Every engine here still cites a wide range of other domains, and your topics may not be ones the dominant source covers at all.
"This ranks the best sources." It counts what appeared, produced by retrieval systems with their own constraints. Frequency of citation is not a quality judgement.
"We should build a forum." A community platform's advantage comes from scale and accumulated content, not from the format.
"The numbers are stable enough to plan on." These are snapshots of systems under active development. The honest planning horizon is months, not years.
The costs of optimising for someone else's skew
The main risk is chasing a number that describes somebody else's query set: restructuring content toward a general-knowledge skew when your own topics are specialised spends real effort against the wrong target. Then format cargo-culting — adding a video because one engine reportedly favours multimodal content is only sensible if the video is genuinely useful. Then neutrality theatre: rewriting commercial content to sound encyclopedic without changing what it contains produces a page that is less persuasive and no more citable. And finally measurement drift, where a team optimising toward a published external figure starts reporting against it and stops measuring its own outcomes.
Branded queries behave differently from general ones, and the source that wins them is usually the brand's own site regardless of the wider skew. Local and geographically specific businesses are a different problem again — local AI search, shopping queries and YMYL topics are kept separate here for exactly that reason. What actually arrives from these surfaces is AI referral traffic statistics, and what never arrives is zero-click search statistics.
What would change this page
Publication of the underlying corpus would change the grades immediately — we would rather have the file than the headline. Our own first-party measurement is the planned replacement, which is why this page is written to be superseded. A significant architectural change at any engine would also change it, since these skews reflect current retrieval designs.
Several outcomes of our own measurement would contradict this page, and all of them get published to the null results registry. The reported skews failing to replicate on our query set. Cross-engine overlap coming in much higher than 11%, which would undercut the main strategic conclusion here. Domain-type structure showing no relationship with citation once topic is controlled. Or run-to-run variance turning out high enough that every share table in this field, ours included, is measuring noise as much as preference.
Open question Four questions stay open. How much does citation share vary by topic within a single engine? How stable is a citation set across repeated runs of the same query? Does overlap between engines rise or fall as they mature? And how much of each skew is explained by crawler access rather than content preference?
Access configuration precedes every content decision on this page: check what you are exposing before you restructure anything. New per-engine measurements ship through the newsletter.
Verification status
Every percentage on this page traces to one vendor-reported corpus that is not independently public, and all are graded Partial following the standard set in the provenance audit. The Ahrefs top-10 overlap figure is the exception, graded Traceable on a published method. Once the Citation Index's own query set is live, this page gets rebuilt against a corpus we collected ourselves and can publish in full. Who is doing the measuring, and on whose money, is on the about page.
Namdev, R. (2026). Most-cited domains in AI search, engine by engine (v2). Retrieved from https://ritiknamdev.com/blog/most-cited-domains-ai-search Published under CC BY 4.0 — reuse freely with attribution.
Once the Citation Index is running, this page will be rebuilt on our own first-party corpus rather than vendor-reported figures.