Statistics · Cross-platform

Most-cited domains in AI search, engine by engine

Which source types dominate citations on ChatGPT, Perplexity and Google AI Overviews — reported per engine rather than pooled, because pooling across engines with different samples is what breaks these tables.

Ritik Namdev Ritik Namdev ·Published September 2026 ·Per-engine, not pooled ·13 min read ·Last verified September 2026
The short version

Each major engine has a distinct centre of gravity: ChatGPT toward Wikipedia and encyclopedic sources, Perplexity toward Reddit, Google AI Overviews toward YouTube and multimodal content. Roughly 11% of domains are reportedly cited by both ChatGPT and Perplexity. All of these figures come from one vendor-reported corpus of a stated 680 million citations whose query mix is not public, so they are graded Partial throughout. Read the table per engine; a pooled "most-cited domains" list is the one thing this data cannot support.

What this page establishes
  • 47.9% of ChatGPT top citations reportedly go to Wikipedia and similarly encyclopedic sources — measured on a vendor corpus of a stated 680M citations, query mix undisclosed. Graded partial.
  • 46.7% of Perplexity top citations reportedly go to Reddit, from the same corpus and the same undisclosed query mix. The two figures are comparable to each other and to nothing else.
  • 23.3% of Google AI Overview citations reportedly favour YouTube and multimodal content — a lower concentration than either of the above, on the same corpus.
  • Only ~11% of domains are reportedly cited by both ChatGPT and Perplexity. That overlap figure, not the source shares, is the finding that changes how work is planned.
  • Every percentage on this page traces to one corpus nobody outside the vendor can inspect. The query set drives a citation-share table more than anything else does, and the query set is exactly what is not published.

Which domains does each engine actually cite?

47.9%

of ChatGPT's top citations reportedly go to Wikipedia and similarly encyclopedic sources — the largest single-source concentration reported for any major engine.

Profound — graded partial · Reported 2026
46.7%

of Perplexity's top citations reportedly go to Reddit, on the same corpus and the same undisclosed query mix.

Profound — graded partial · Reported 2026
23.3%

of Google AI Overview citations reportedly favour YouTube and multimodal content — a notably lower concentration than the other two.

Profound — graded partial · Reported 2026
EngineReported dominant source typePopulation behind the figureStructural implication for a publisherGrade
ChatGPTWikipedia and encyclopedic reference — 47.9% of top citationsVendor corpus, stated 680M citations; query mix, date range and counting rule undisclosedNeutral, comprehensive, non-promotional coveragePartial
PerplexityReddit and community discussion — 46.7% of top citationsSame corpus, same undisclosed query mixSpecific, first-hand, frequently updated detailPartial
Google AI OverviewsYouTube and multimodal — 23.3% of citationsSame corpus; concentration roughly half that of the other twoA genuine visual or video companion, not a token onePartial
ChatGPT (index lineage)A large majority of SearchGPT citations matched Bing's top resultsSeer Interactive's own sample; method described, data not publishedBing visibility is a load-bearing input most teams never checkPartial
Cross-engine (vs. classic Google)About 12% of AI-cited URLs rank in Google's top 10Ahrefs, large URL sample; method described publiclyClassic rank is a weak predictor of citationTraceable
ClaudeNo comparable per-engine share published—Treat any Claude source-mix claim as unmeasuredBroken chain

Per-engine pages carry the detail: ChatGPT, Perplexity, Google AI Overviews, Claude and Gemini. What each dominant source stands in for has its own studies: the Wikipedia dependency study asks what ChatGPT cites when the encyclopedia has nothing on a topic, and the Reddit dependency study asks the same of Perplexity. Both are gap studies, and the gaps are where an ordinary publisher can actually win.

Why a single pooled most-cited list is the wrong asset

The genre convention is a ranked list of domains under a heading like "the 50 most-cited sites in AI search." That list cannot be built honestly from this data. Building it anyway is the weakness in almost every version you will find.

Pooling requires that each engine contributed a comparable sample. Nothing guarantees that here. The engines surface different numbers of sources per answer, so an engine that shows fifteen citations contributes more rows than one that shows three, from identical retrieval behaviour. The counting rule — per answer, per query, or per domain — is not stated, and each produces a different number from the same data. And a pooled share is a property of the query set as much as of the engines: change the questions and the table changes with no engine behaviour altered at all.

A pooled figure therefore describes a population that does not exist — a composite engine nobody uses. The per-engine rows above describe real, if unverifiable, populations. That is why the table is the asset on this page and the ranked list is not.

ChatGPT leans on Wikipedia. Perplexity leans on Reddit. Google AI Overviews lean on YouTube. Only ~11% of domains are cited by both ChatGPT and Perplexity. 'AI visibility' isn't one target — it's at least three.

Share on X

How much do the engines agree? About 11%

Domain overlap between ChatGPT and Perplexity citations
  • Cited by both ChatGPT and Perplexity 11
  • Cited by only one 89
Source: Profound, reported figure — graded partial, underlying corpus not public.

An ~11% overlap means optimising for one engine's citation pattern gives you very little assurance of appearing in another's. This is the strongest available argument for measuring visibility per platform rather than as one blended score, consistent with the measurement standard.

Read it carefully, though. It says most domains cited by one engine are not cited by the other. It does not say the engines disagree about facts, or that one is better. It is not a probability for your specific site — a site with strong coverage of a topic may well be cited everywhere, and a weak one nowhere. Overlap also depends on how many domains each engine cites in total. If one draws from a much wider pool, low overlap is partly a consequence of pool size rather than of disagreement, and the pool sizes are not published either.

Overlap with classic Google rankings is a separate question, measured by Ahrefs across a large URL sample at about 12% of AI-cited URLs ranking in the top 10. Overlap between engines is the subject of the cross-platform citation concordance study, which is the work this section is really waiting on.

47.9%

of ChatGPT's top citations reportedly go to Wikipedia and similarly encyclopedic sources — the largest single-source concentration reported for any major engine.

Profound · graded partial

Why domain type, not domain authority, predicts citation

The pattern is best explained by retrieval architecture, not by any property of the cited domains. ChatGPT Search has historically leaned on a Bing-derived index with a bias toward broad, structured reference content. Seer Interactive found a large majority of SearchGPT citations matched Bing's top results, which makes the Microsoft index load-bearing for a product most people never associate with it. The Bing and Copilot guide covers that surface directly. Perplexity's own crawler and citation-first interface are associated with community-vetted, frequently updated discussion. Google's AI Overviews can draw on YouTube and the Knowledge Graph, inputs the other two do not have natively.

None of this is primarily about backlink profiles. The claim that backlink-derived authority matters less than brand presence is examined in brand mentions versus backlinks, with Zyppy's correlational work among the few public attempts to quantify any of it.

Two structural factors get asserted far more often than tested. First, structured data, where Ahrefs found no clear effect on its own corpus and our schema study takes it up; and named, credentialed authorship, the subject of the author E-E-A-T study. Every tactic in that family is graded in the tactic evidence scoreboard.

A skew can also arise for four different reasons that look identical from outside, and they imply different actions.

MechanismWhat it meansCan a publisher act on it?
Index compositionThe engine's underlying index simply contains more of that source type.Only by being in that index at all.
Retrieval weightingThe ranking step favours certain structural or authority signals.Partly — structure is controllable.
AvailabilityThe source is openly licensed, unpaywalled and easy to fetch.Yes — access is a choice, and the least discussed one.
Query mixThe measured questions happen to suit that source type.Not at all — this is a measurement artefact.

The availability row is the most actionable. A source that is blocked or hard to fetch cannot be cited whatever its quality, and that is a configuration decision rather than a content one. Each operator publishes its own agents — OpenAI, Perplexity and Anthropic — consolidated in the bot user-agent registry. Whether those crawlers render JavaScript decides whether your content exists to them at all, and the technical audit is the checklist version. Cloudflare's report on undeclared crawlers is a reminder that the declared list and the actual traffic are different objects.

Why every number here is graded partial

A partial grade means the figure traces to a named source that describes its own data, but the data itself cannot be inspected by anyone outside that source. We can name who said it. We cannot check it.

Four things specifically are missing, and each of them moves the number. The query set, which drives everything: a corpus weighted toward general-knowledge questions shows encyclopedic dominance almost by construction, and one weighted toward product research does not. The date range, since retrieval systems change without announcement. The country and language mix. And the definition of a "top citation" — counted once per answer, once per query, or once per domain, which changes the arithmetic substantially.

None of this makes the figures wrong. Unverifiable is a distinct category from wrong. The direction here is corroborated in independent write-ups rather than independent data: Discovered Labs on how each assistant chooses sources, Leapd's cross-platform summary and Ziptie on ChatGPT's source selection all describe the same lean without re-measuring it. Corroboration and replication are different things. We report the direction confidently and the decimals cautiously.

What bends a most-cited-domains result

If you are evaluating someone else's measurement, or running your own, these are the levers. Query set composition dominates everything else. Personalisation and location mean a corpus collected in one region describes that region. Time makes every figure a snapshot — how quickly a new page enters the citable pool is the crawl-to-citation latency question, and how long a citation survives is the half-life question.

Counting rules and deduplication of subdomains and regional variants both change concentration figures noticeably. And sampling of answers matters: engines can answer the same question differently on repeated attempts, and no published figure we have seen reports that variance.

Building your own citation sample

A small first-party sample beats a large corpus you cannot inspect, provided you are honest about its limits.

Building a first-party citation sample
  1. 01 Freeze a question list Thirty to fifty real customer questions, in natural language
  2. 02 Fix the counting rule Decide per-answer, per-query or per-domain before you start
  3. 03 Run every engine On at least three separate days, because single samples hide variance
  4. 04 Record everything Including refusals and answers with no citations at all
  5. 05 Tabulate by type The domain-type table is more instructive than the domain table
  6. 06 Publish the method To yourself at minimum — otherwise it is a snapshot, not a measurement

Freeze thirty to fifty questions a real customer would ask, in natural language rather than keywords. Decide the counting rule before you start — we suggest recording every cited domain once per answer and keeping the raw answer text. Run each question on each engine you care about, across at least three separate days.

Record everything, including questions that produced no citations and answers that refused; dropping those biases the result. Tabulate by domain and by domain type, because the type table tells you what to build rather than who to envy. Then write down the method — the query list, the dates, the rule. If you cannot repeat the exercise in three months and get a comparable number, you built a snapshot rather than a measurement. The zero-to-cited log study is our own worked version of this.

Common misreadings of a citation-share table

One site takes half, so we are locked outA plurality is not a monopoly
This ranks the best sourcesFrequency of citation is not a quality judgement
So we should build a forumThe advantage is scale and accumulation, not format
The numbers are stable enough to plan onSnapshots of systems under active development

"Half of citations go to one site, so the rest of us are locked out." A plurality is not a monopoly. Every engine here still cites a wide range of other domains, and your topics may not be ones the dominant source covers at all.

"This ranks the best sources." It counts what appeared, produced by retrieval systems with their own constraints. Frequency of citation is not a quality judgement.

"We should build a forum." A community platform's advantage comes from scale and accumulated content, not from the format.

"The numbers are stable enough to plan on." These are snapshots of systems under active development. The honest planning horizon is months, not years.

The costs of optimising for someone else's skew

The main risk is chasing a number that describes somebody else's query set: restructuring content toward a general-knowledge skew when your own topics are specialised spends real effort against the wrong target. Then format cargo-culting — adding a video because one engine reportedly favours multimodal content is only sensible if the video is genuinely useful. Then neutrality theatre: rewriting commercial content to sound encyclopedic without changing what it contains produces a page that is less persuasive and no more citable. And finally measurement drift, where a team optimising toward a published external figure starts reporting against it and stops measuring its own outcomes.

Branded queries behave differently from general ones, and the source that wins them is usually the brand's own site regardless of the wider skew. Local and geographically specific businesses are a different problem again — local AI search, shopping queries and YMYL topics are kept separate here for exactly that reason. What actually arrives from these surfaces is AI referral traffic statistics, and what never arrives is zero-click search statistics.

What would change this page

Publication of the underlying corpus would change the grades immediately — we would rather have the file than the headline. Our own first-party measurement is the planned replacement, which is why this page is written to be superseded. A significant architectural change at any engine would also change it, since these skews reflect current retrieval designs.

Several outcomes of our own measurement would contradict this page, and all of them get published to the null results registry. The reported skews failing to replicate on our query set. Cross-engine overlap coming in much higher than 11%, which would undercut the main strategic conclusion here. Domain-type structure showing no relationship with citation once topic is controlled. Or run-to-run variance turning out high enough that every share table in this field, ours included, is measuring noise as much as preference.

Open question Four questions stay open. How much does citation share vary by topic within a single engine? How stable is a citation set across repeated runs of the same query? Does overlap between engines rise or fall as they mature? And how much of each skew is explained by crawler access rather than content preference?

Next step

Access configuration precedes every content decision on this page: check what you are exposing before you restructure anything. New per-engine measurements ship through the newsletter.

Verification status

Every percentage on this page traces to one vendor-reported corpus that is not independently public, and all are graded Partial following the standard set in the provenance audit. The Ahrefs top-10 overlap figure is the exception, graded Traceable on a published method. Once the Citation Index's own query set is live, this page gets rebuilt against a corpus we collected ourselves and can publish in full. Who is doing the measuring, and on whose money, is on the about page.

How to cite this
Namdev, R. (2026). Most-cited domains in AI search, engine by engine (v2). Retrieved from https://ritiknamdev.com/blog/most-cited-domains-ai-search

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Once the Citation Index is running, this page will be rebuilt on our own first-party corpus rather than vendor-reported figures.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Discovered Labs — how ChatGPT, Claude, Perplexity and AIO cite sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Ahrefs — only 12% of AI-cited URLs rank in Google's top 10ahrefs.com/blog/ai-search-overlap Discovered Labs — how ChatGPT, Claude and Perplexity choose sourcesdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Leapd — how ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Ziptie — how ChatGPT chooses its sourcesziptie.dev/blog/how-does-chatgpt-choose-its-sources Ziptie — how original research wins AI citationsziptie.dev/blog/how-original-research-wins-ai-citations Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Perplexity — official crawler documentationdocs.perplexity.ai/docs/resources/perplexity-crawlers OpenAI — official bot and crawler documentationplatform.openai.com/docs/bots Anthropic Support — does Anthropic crawl the web, and how to block itsupport.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Cloudflare — from Googlebot to GPTBot: who is crawling your siteblog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Cloudflare — Perplexity using stealth, undeclared crawlersblog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives BuzzStream — study of publishers blocking AI crawlerswww.buzzstream.com/blog/publishers-block-ai-study Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors Onely — writing LLM-friendly contentwww.onely.com/blog/llm-friendly-content Ahrefs — no clear schema effect on AI citationsahrefs.com/blog/schema-ai-citations Wikipedia — Generative engine optimizationen.wikipedia.org/wiki/Generative_engine_optimization GEO — original arXiv paper on generative engine optimizationarxiv.org/abs/2311.09735 Statista — global monthly ChatGPT userswww.statista.com/statistics/1659718/global-monthly-chatgpt-users
FAQ

Frequently asked questions

Does this mean I should focus on getting mentioned on Reddit or Wikipedia specifically?
Not directly for most brands. Those platforms dominate because of what they are: large, structured, community-vetted knowledge bases. It is not because any individual page there is easy to get cited. The transferable lesson is about content type and structure, not "go post on Reddit."
Why don't engines agree more on what to cite?
Each engine has a different retrieval architecture, index, and citation-selection logic. ChatGPT leans on encyclopedic sourcing. Perplexity leans on community and forum content. Google AI Overviews lean on their own multimodal index. They are solving the same problem with genuinely different tools.
Where does the 680-million-citation figure come from?
A vendor-reported corpus size, from Profound, not independently verified. It is graded partial on our provenance scale. It is a plausible scale given the vendor's stated reach, but the underlying data is not public.
If my brand cannot become Wikipedia or Reddit, is ChatGPT or Perplexity citation simply unreachable?
No. The skew describes the dominant category, not an exclusive one. Brand and publisher domains still make up a meaningful share of citations on every engine. The skew explains what wins the plurality, not what wins everything.
Can I reproduce these percentages myself?
Not exactly, and that is the core problem with them. The underlying corpus is not public, so nobody outside the vendor can recompute the figures or check the query mix behind them. What you can do is build your own smaller sample for your own topics, which will not match theirs but will describe your situation.
Is domain authority (DR/DA) still relevant to AI citation?
It correlates weakly at best in available data. Brand mentions and content structure appear to be more strongly associated with citation than backlink-based authority metrics, though causal evidence for any of these factors remains thin. See the GEO tactic scoreboard for the full grading.
How often would this measurement need repeating to stay useful?
Often enough to catch retrieval changes, which happen without announcement. We would treat any figure older than a few months as historical rather than current. Repeated measurement is only meaningful if the method stays identical, and most published figures do not say enough about their method for anyone to repeat it at all.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.