Original research · Pre-registered

The ChatGPT Wikipedia Dependency

ChatGPT's citations reportedly skew toward Wikipedia at nearly 48%. For the large share of topics Wikipedia doesn't cover well, what fills that gap? Nobody has published an answer.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — design stage ·14 min read ·Last verified September 2026
The short version

Reported data puts Wikipedia at nearly 48% of ChatGPT's top citations. That statistic describes topics where Wikipedia exists and is comprehensive. For the large share of topics where it doesn't — niche products, emerging concepts, specific technical questions — what replaces it? This page registers a study to find out, and sits inside the same programme as the AI citation index and the parallel Reddit-dependency study on Perplexity.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. The gap-topic query set has not been built and no citations have been collected.

What is known
  • ChatGPT's widely-quoted ~48% Wikipedia citation share describes topics Wikipedia covers well. Its provenance is traced in full on this page.
  • Wikipedia's structural properties — broad coverage, neutral tone, stable URLs — are observable and shared by few other sources.
What is not yet known
  • What wins the citation on the much larger set of topics Wikipedia does not cover.
  • Whether citation concentration is lower on gap topics than on covered ones.
  • Whether the skew reflects a hard-coded preference or the properties Wikipedia happens to have.

The question nobody has answered

Every published Wikipedia-skew statistic implicitly describes a subset of queries: ones where Wikipedia has relevant, substantial coverage. It says nothing about the queries where that isn't true. Outside general encyclopedic knowledge, that's most of them.

Say a content creator's topic genuinely has no strong Wikipedia article. Knowing what wins the citation in that gap is more useful to them than knowing Wikipedia's overall dominance.

Scale is what makes it matter. ChatGPT's monthly user numbers, tracked by Statista and cross-checked against Similarweb, put it well ahead of the other answer engines. A source-selection quirk in a small tool is trivia; the same quirk in the most-used answer engine is infrastructure.

What is known

Evidence

ChatGPT's citations reportedly skew toward Wikipedia and encyclopedic sources at roughly 47.9% of top citations — Partial , vendor-reported from a non-public corpus.

Open question

What specifically wins the citation for topics without meaningful Wikipedia coverage — no published study addresses this directly.

47.9%

of ChatGPT's top citations reportedly go to Wikipedia and encyclopedic sources. Vendor-reported from a corpus nobody outside the vendor can inspect.

Profound

Independent write-ups point the same way without settling the number. Ziptie on how ChatGPT chooses its sources and Discovered Labs on citation patterns across ChatGPT, Claude and Perplexity both describe an encyclopedic pull on general-knowledge queries. Neither publishes a corpus, which is why every figure on this site carries a verification grade beside it.

Why Wikipedia wins when it's available

Four properties make Wikipedia the default. Broad, reasonably neutral coverage. A predictable structure of definition, history, key facts, references. Frequent updates. And sheer scale — an article already exists for an enormous share of general-knowledge queries.

Any source filling the gap would plausibly need some subset of those properties: structure and reasonable comprehensiveness, if not scale. That is why the hypotheses below predict official documentation and established publications, rather than arbitrary sources, as the likely runners-up.

The prediction sits close to what the original GEO paper argued about how generative systems select sources, and to Onely's account of what makes content LLM-friendly.

Defining "thin coverage" precisely

To avoid a post-hoc, cherry-picked definition of "gap topic," this study specifies a concrete threshold before data collection begins. A topic qualifies as thin-coverage if no Wikipedia article exists for it directly. It also qualifies if the closest matching article mentions the topic only in passing. A single sentence or list entry within a much broader article doesn't count as a dedicated subject.

This threshold gets published up front. Describing it only after seeing which topics produced interesting results wouldn't keep the finding honest. Publishing it first is what keeps this a genuine test, not a selectively illustrated one.

The actual research question

Restrict the analysis to queries about topics with thin or no Wikipedia coverage. Which domain types, structural properties, and content formats capture the citation instead? Does the "runner-up" source resemble Wikipedia's properties, neutral tone, broad structure, comprehensive coverage, or does it look like something else entirely?

PropertyWikipedia-covered topicGap topic
Default sourceA dedicated encyclopedia article exists and is usually currentNone; the slot is contested by whatever else covers the subject
Likely competitor setEffectively one dominant source plus supporting citationsDocumentation, trade publications, retailers, forums, in unknown proportion
Neutrality of the winnerEditorially enforced, non-commercial by policyFrequently commercial, with no neutrality requirement at all
Realistic strategyCompete on adjacent questions rather than the definitional oneBuild the reference page an encyclopedia would have written
Evidence availableVendor-reported citation-share figures, ungraded corporaNone published; this is the hole the study is registered to fill

Study design

Collection pipeline
  1. 01 Identify topics Where Wikipedia has no dedicated article
  2. 02 Query ChatGPT Across the fixed query set for those topics
  3. 03 Log citations What replaces the encyclopedic default
  4. 04 Classify sources By domain type and structure
  5. 05 Compare Against topics where Wikipedia does exist
Study stages and what gets published at each
  1. v0Design stage

    Threshold, query set and hypotheses published before collection

    Registering the definition first is what keeps this a test rather than an illustration.

  2. v2Classification

    Cited domains coded by type and structural properties

    Two coders, with disagreement rates reported.

  3. v3Publication

    Results, nulls and raw files released together

    A contrary result is published with the same prominence as a confirmation.

Two design choices matter most, because they are where studies like this usually go wrong. The query set is fixed before collection and repeated on a schedule — single-run answers are unstable, and a one-shot sample finds whatever the sampler hoped for. The classification scheme for cited domains is written down in advance, the same discipline the whole Index programme works under.

Pre-registered hypotheses

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
WD1In the absence of Wikipedia coverage, official documentation and established industry publications capture a disproportionate share of citationsExpect to hold
WD2Citation concentration is lower (more distinct domains cited) for gap-topic queries than for Wikipedia-covered onesExpect to hold
WD3The runner-up source's structural properties (broad coverage, neutral tone) more closely resemble Wikipedia than a randomly selected competing source wouldExpect to hold

ChatGPT's ~48% Wikipedia citation share describes topics Wikipedia actually covers well. Nobody has published what wins the citation for the much larger set of topics it doesn't.

Share on X

One more prediction is worth recording, because it is cheap to check. Suppose gap citations go to whatever most resembles an encyclopedia entry. The winning pages should then look like the ones described in how to get cited by ChatGPT — direct definitions, self-contained facts, clean headings — rather than simply the pages that rank well. Ahrefs' analysis of which pages AI Overviews cite found the overlap with top-ranking pages weaker than expected, which is consistent with structure mattering independently of position.

Why an encyclopedia is structurally easy to cite

The skew is often described as a preference. It is more useful to read it as five ordinary properties stacking up, none of which requires deliberate favouring.

The five properties this page argues account for the encyclopedic skew, and what each one does for a retrieval system. A description of a mechanism, not a measurement of one — none of these has been isolated experimentally.
PropertyWhy it favours citationWhere alternatives fail
Coverage breadthA source that exists for nearly every topic wins by availability aloneMost reference sites cover one domain — which is exactly the gap this study targets
Predictable structureDefinition, then history, then detail is close to ideal for a system extracting a short factual claim; it does not have to huntCommercial pages bury the definition below a pitch
Licence and accessibilityOpenly licensed, unpaywalled, technically simple to fetchMany high-quality alternatives fail on at least one of the three
Neutral registerProse that sells nothing is safer to quote than a commercial page making the same claimNo neutrality requirement applies to the sources contesting a gap topic
Link densityAmong the most linked-to domains on the web, so any retrieval system using authority signals surfaces itA new reference page starts with none of it

Link density deserves a caveat. Wikipedia is among the most linked-to domains on the web, so any retrieval system using authority signals surfaces it. Whether links or unlinked prominence carry more of that weight is genuinely contested. The correlational case is set out in brand mentions versus backlinks, and all of that evidence is observational.

Hypothesis We expect these properties, not a hard-coded preference, explain most of the skew. One caveat: "this property looks like it should matter" has a poor track record here. The schema markup study is registered precisely because that reasoning failed for llms.txt. The testable consequence is that gap-topic citations should go to whichever source best approximates the same stack — hypothesis WD3.

Where the 47.9% figure actually comes from

A number gets stated with its provenance or not at all, so: the headline figure is vendor-reported and the corpus behind it is not public. We cannot check the query set, the sampling dates, the country, or how a "top citation" was defined. That is why it carries a Partial grade rather than a stronger one.

We use it anyway, deliberately: the direction is corroborated by ordinary observation — ask a general-knowledge question and encyclopedic sources appear constantly. It is the decimal we do not trust, not the pattern.

One structural detail deserves more attention than the decimal. ChatGPT's web results have historically leaned on Bing's index, and Seer Interactive found that 87% of SearchGPT citations matched Bing's top results. If that relationship holds, part of the encyclopedic skew is inherited from a conventional search index rather than chosen by the model. That is a different claim with different implications.

None of the gap-topic design depends on the figure being 47.9 rather than 40 or 55. The study measures its own baseline from its own query set, so a badly wrong vendor figure leaves it standing. That independence was built in deliberately.

The circularity problem nobody wants to discuss

One structural issue underneath all of this deserves stating plainly, even though this study cannot resolve it. Wikipedia articles are built from citations to other sources. If AI answers cite Wikipedia heavily, and people increasingly get their facts from AI answers, then a large share of public knowledge flows through one editorial pipeline. That pipeline has known coverage biases: by language, by geography, by topic, and by who volunteers to edit.

A second-order version is worse. If AI-generated text spreads across the web, and Wikipedia editors later cite web sources that were themselves AI-generated from Wikipedia, the loop closes. A claim could eventually be supported only by its own echo.

Open question We have no measurement of how much of this is happening, and we are not going to speculate with numbers. It sits next to another unmeasured question: how long a cited page stays cited, the subject of the citation half-life study. It matters most where being wrong is expensive, which is why YMYL content in AI search is tracked separately. The mechanism is not exotic and the incentives point toward it growing — and no citation-share statistic captures any of it.

Confounds in a gap-topic study

Suppose the study finds that documentation and established publications fill the gap. Before believing WD1, several alternative explanations have to be ruled out.

ConfoundWhy it could produce the same result
Topic-age imbalanceGap topics skew newer. Newer topics have different source ecosystems regardless of Wikipedia's absence.
Commercial densityMany gap topics are commercial, where documentation and retailer content simply dominate the available web.
Query phrasingGap-topic queries may be phrased more specifically, which favours documentation over general reference by lexical match alone.
Wikipedia's own quality gradientTopics with no article often also have thinner coverage everywhere. Scarcity, not substitution, may drive the result.
Selection of the gap setIf the thin-coverage threshold is applied loosely, the set fills with topics chosen for interestingness rather than by rule.

The first two are the serious ones. Both would produce exactly the predicted pattern with no substitution behaviour involved at all. Matching gap and non-gap topics on age and commercial intent is therefore not optional. It is what makes the comparison mean anything.

What does not transfer to other engines

Everything on this page is about one engine, deliberately. Different engines draw on different source pools and different retrieval architectures. A skew toward community forums on one engine and toward encyclopedic sources on another are not two views of one phenomenon; they are different systems making different choices. Per-engine citation shares differ enough that a blended "AI search" figure hides more than it shows. Cross-platform citation concordance is the protocol registered to measure how small the overlap really is, and Ahrefs reached a similar conclusion measuring overlap between AI search engines.

Gap behaviour is even less transferable than the baseline: what fills a hole depends on what is nearby in that engine's source pool. An engine leaning on community content fills a gap with community content; one leaning on a classic web index fills it with whatever that index ranks.

Within Google alone the point holds. AI Mode and AI Overviews are not one system, and the source delta between them has its own registered protocol.

Hypothesis We expect gap-fill behaviour to vary more between engines than baseline behaviour does, because gaps expose the underlying source pool rather than the shared default. That is a prediction the parallel study on a second engine is designed to test.

How to check your own topic for a Wikipedia gap

You can do a rough version of this analysis for your own topic set in an hour. It will not be publishable research, but it will tell you which strategy applies to you.

  1. List your twenty most commercially important topics as plain-language subjects, not keywords.
  2. For each, check whether a dedicated Wikipedia article exists, applying the same threshold this study uses: a passing mention inside a broader article does not count.
  3. Sort into three buckets — strong article, thin or passing mention, nothing at all.
  4. For the gap buckets, ask an AI engine the natural question a reader would ask and record which sources it names. Repeat on a different day; single samples are unreliable.
  5. Look at what the cited sources have in common structurally, not just who they are — broad reference pages, documentation, comparison articles. That shape is your target.

Before any of it, confirm the engine can reach you at all. Most AI crawlers do not execute JavaScript, so a reference page whose text arrives client-side is invisible however encyclopedic it reads. Check that retrieval bots are allowed through in robots.txt, then run the technical GEO audit over the page itself.

Two categories need separate treatment. Local queries resolve against a different source pool entirely, and shopping queries pull retailer and review content no encyclopedia was going to supply. Treat the whole exercise as reconnaissance — no control group, no repeated sampling, no statistical power — which is still more than most content planning starts with.

Null results we would publish

Four outcomes would be null or contrary, and each gets published with the same prominence as a confirmation, in the null results registry:

  • Gap-topic citations show no coherent structural pattern at all — which would mean the advice to build encyclopedic reference pages has no support from this line of evidence, worth knowing before anyone spends a year on it.
  • Citation concentration is the same or higher for gap topics, contradicting WD2.
  • The runner-up source resembles Wikipedia no more than a random competitor does, contradicting WD3.
  • A definitional failure: a thin-coverage threshold two people cannot apply consistently to the same topic list. Boring, and genuinely useful — it would tell everyone else attempting this study where the work actually is.

Open questions

Open question How much of a gap-fill citation is explained by page structure versus domain authority? The two are almost always confounded in observation.

Open question Does the encyclopedic skew differ by language, and how badly does it hurt topics whose best sources are not in English?

Open question When Wikipedia coverage for a topic improves, does citation shift toward it, and how quickly? A natural experiment nobody has run.

Open question How much of the web that engines cite is already downstream of Wikipedia? The circularity question, restated as something measurable in principle.

This study is one slice of the AI Citation Index programme. The crawl-to-citation latency study answers the timing half of the gap question: how long after publishing a reference page you could expect to see anything.

When this reports

No gap-topic queries have been run. The results, the topic list and the raw citation records go out to the newsletter when the first collection window closes. To propose a gap topic for the set, or to point out a flaw in the design before it runs, how to reach me is on the about page — there is no sign-up form, just a conversation.

Limitations

  • "Thin Wikipedia coverage" requires a defined threshold, which will be specified and published before data collection to prevent post-hoc redefinition.
  • Results are specific to ChatGPT — the same question for Perplexity's Reddit dependency is registered as a separate, parallel study.
  • The baseline figure is vendor-reported, from a corpus nobody outside the vendor can inspect, and is used here as a direction rather than a measurement.
  • Gap and non-gap topics must be matched on age and commercial intent, or the two most serious confounds above reproduce the predicted result with no substitution behaviour involved.
  • English-language, single-region collection. Wikipedia's coverage depth varies enormously by language, so a gap in the English encyclopedia is not a gap everywhere, and no figure here should be read as describing non-English retrieval.
  • Topic selection is a judgement call. The thin-coverage threshold is published in advance, but which twenty verticals the gap topics are drawn from will shape the runner-up mix. The topic list ships with the data so the choice can be argued with.
  • Measurement is of what the answer surface displays, not of what the model retrieved. A source consulted but not named leaves no trace, so citation share is a floor on source use, not a full account of it.
  • Reproduction is approximate. Anyone can re-run the published query set, but a later run measures a later product state and a later encyclopedia; agreement on direction, not on the decimal, is the realistic bar.
How to cite this
Namdev, R. (2026). The ChatGPT Wikipedia Dependency (v1). Retrieved from https://ritiknamdev.com/blog/wikipedia-dependency-chatgpt

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Builds directly on the source-skew finding in ChatGPT citation statistics and most-cited domains in AI search.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Ziptie — How does ChatGPT choose its sources?ziptie.dev/blog/how-does-chatgpt-choose-its-sources Ziptie — How original research wins AI citationsziptie.dev/blog/how-original-research-wins-ai-citations Discovered Labs — AI citation patterns across ChatGPT, Claude and Perplexitydiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Discovered Labs — How each platform cites sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Leapd — How ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Ahrefs — Overlap between AI search enginesahrefs.com/blog/ai-search-overlap Ahrefs — Which pages AI Overviews cite (top-10 analysis)ahrefs.com/blog/ai-overview-citations-top-10 Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors RankScience — AI citations, brand mentions and the visibility gapwww.rankscience.com/blog/ai-citations-brand-mentions-visibility-gap Aggarwal et al. — GEO: Generative Engine Optimization (arXiv)arxiv.org/abs/2311.09735 Wikipedia — Generative engine optimizationen.wikipedia.org/wiki/Generative_engine_optimization Statista — Global monthly ChatGPT userswww.statista.com/statistics/1659718/global-monthly-chatgpt-users Similarweb — Generative AI usage statisticsaisearch.similarweb.com/blog/gen-ai-stats Semrush — AI Overviews studywww.semrush.com/blog/semrush-ai-overviews-study Salespeak — Content freshness and AI searchsalespeak.ai/aeo-news/content-freshness-ai-search Onely — What makes content LLM-friendlywww.onely.com/blog/llm-friendly-content Anthropic — Web search tool documentationplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool OpenAI — Bots and crawler documentationplatform.openai.com/docs/bots Margen — Perplexity statistics 2026www.margen.net/perplexity-statistics-2026 StatCounter — Search engine market sharegs.statcounter.com/search-engine-market-share
FAQ

Frequently asked questions

Why focus on ChatGPT specifically for this question?
Because it shows the strongest reported Wikipedia skew of the major engines, roughly 47.9% of top citations in vendor-reported data. That makes the "what happens when the default isn't available" question most consequential there.
Isn't it obvious that something else gets cited when Wikipedia has nothing?
Obviously something fills the gap. The interesting question is what kind of source. Does it resemble Wikipedia's structural properties, broad, neutral, well-organized, or does it look completely different? That pattern would tell content creators what to emulate for topics an encyclopedia doesn't cover well.
How common are topics without Wikipedia coverage?
Very common, outside general-knowledge domains. Most product categories, niche technical topics, and emerging concepts have thin or no Wikipedia coverage. That's exactly why this question has real practical value for a large share of content.
Would a company be better off just trying to get a Wikipedia page created?
For many commercial topics, no. Wikipedia's notability guidelines and neutral-point-of-view policy make it a poor fit for product pages, brand pages, and most commercial content. That's a large part of why this gap exists structurally, not by oversight. The more actionable path for most brands is winning the gap-fill citation this study is designed to characterize, not competing to become the encyclopedic default itself.
How firm is the 47.9% figure itself?
It is vendor-reported from a corpus nobody outside the vendor can inspect, which is why it is graded partial rather than traceable. We cite it because it is the best available signal of a real pattern, not because it is verified. If the underlying corpus were published, the figure might move considerably in either direction.
Does Wikipedia dominance mean AI answers are more reliable?
Not on its own, and treating it that way would be a mistake. Wikipedia quality varies enormously by article: heavily-edited general topics are usually solid, obscure ones can sit unreviewed for years. A high citation share tells you about source selection, not about whether the resulting answer was correct.
Should I try to get my brand mentioned inside an existing Wikipedia article?
Only if the mention genuinely belongs there by Wikipedia's own standards, and adding it yourself is against their conflict-of-interest guidance. Editing an encyclopedia for citation advantage is a bad-faith use of a public resource, and it usually gets reverted. The gap strategy on this page exists precisely because it does not require touching Wikipedia at all.
What if a topic has a Wikipedia article that is simply bad?
That is a third category the study design has to handle, and it sits between the two we defined. A thin but existing article may still capture the citation, or may be passed over for a better source. We will record article quality signals alongside existence, and report those cases separately rather than forcing them into a binary.
Could this study's findings vary a lot by topic category?
Plausibly. A gap in coverage of an emerging technology concept, and a gap in coverage of a niche consumer product, likely get filled by different kinds of sources: technical documentation and forums, versus retailer and review sites. Breaking results down by topic category, not just reporting one blended finding, is part of the planned analysis.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.