Original research · Pre-registered

The ChatGPT Wikipedia Dependency

ChatGPT's citations reportedly skew toward Wikipedia at nearly 48%. For the large share of topics Wikipedia doesn't cover well, what fills that gap? Nobody has published an answer.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — design stage ·11 min read
The short version

Reported data puts Wikipedia at nearly 48% of ChatGPT's top citations. That statistic describes topics where Wikipedia exists and is comprehensive. For the large share of topics where it doesn't — niche products, emerging concepts, specific technical questions — what replaces it? This page registers a study to find out.

The question nobody has answered

Every published Wikipedia-skew statistic implicitly describes a subset of queries: ones where Wikipedia has relevant, substantial coverage. It says nothing about the queries where that isn't true. Outside general encyclopedic knowledge, that's most of them.

Say a content creator's topic genuinely has no strong Wikipedia article. Knowing what wins the citation in that gap is more useful to them than knowing Wikipedia's overall dominance.

What is known

Evidence

ChatGPT's citations reportedly skew toward Wikipedia and encyclopedic sources at roughly 47.9% of top citations — Partial , vendor-reported from a non-public corpus.

Open question

What specifically wins the citation for topics without meaningful Wikipedia coverage — no published study addresses this directly.

Why Wikipedia wins when it's available

Understanding the gap requires first naming what makes Wikipedia the default in the first place: broad, reasonably neutral coverage. A consistent, predictable structure: definition, history, key facts, references. Frequent updates. And, not least, sheer scale. Wikipedia simply has an article for an enormous share of general-knowledge queries already.

Any source expected to fill the gap where Wikipedia has nothing would plausibly need to replicate some subset of these properties, structure and reasonable comprehensiveness, if not scale. That's one of the reasons this study's hypotheses, below, predict official documentation and established publications, rather than arbitrary sources, as the more likely runners-up.

Defining "thin coverage" precisely

To avoid a post-hoc, cherry-picked definition of "gap topic," this study specifies a concrete threshold before data collection begins. A topic qualifies as thin-coverage if no Wikipedia article exists for it directly. It also qualifies if the closest matching article mentions the topic only in passing. A single sentence or list entry within a much broader article doesn't count as a dedicated subject.

This threshold gets published up front. Describing it only after seeing which topics produced interesting results wouldn't keep the finding honest. Publishing it first is what keeps this a genuine test, not a selectively illustrated one.

The actual research question

Restrict the analysis to queries about topics with thin or no Wikipedia coverage. Which domain types, structural properties, and content formats capture the citation instead? Does the "runner-up" source resemble Wikipedia's properties, neutral tone, broad structure, comprehensive coverage, or does it look like something else entirely?

Study design

Collection pipeline
  1. 01 Identify topics Where Wikipedia has no dedicated article
  2. 02 Query ChatGPT Across the fixed query set for those topics
  3. 03 Log citations What replaces the encyclopedic default
  4. 04 Classify sources By domain type and structure
  5. 05 Compare Against topics where Wikipedia does exist

A worked hypothetical gap-topic example

Here's what the study is looking for, using an invented scenario. Consider a query about a specific, relatively new category of home-fitness equipment, with no dedicated Wikipedia article, only a brief passing mention inside a broader "fitness equipment" article. Under the study's design, a query about this category would qualify as a gap topic.

The prediction under hypothesis WD1: the citation in this case is more likely captured by a source like an established buying-guide publication, or a manufacturer's own comprehensive documentation page. These sources aren't Wikipedia. But they still offer the kind of broad, structured coverage Wikipedia would have provided, had an article existed. A single narrow product listing or an unrelated forum thread is less likely to win. Does real data actually bear this out, across many such gap topics, not just one invented example? That's exactly what the study is designed to determine.

Pre-registered hypotheses

#HypothesisPrediction
WD1In the absence of Wikipedia coverage, official documentation and established industry publications capture a disproportionate share of citationsSupported
WD2Citation concentration is lower (more distinct domains cited) for gap-topic queries than for Wikipedia-covered onesSupported
WD3The runner-up source's structural properties (broad coverage, neutral tone) more closely resemble Wikipedia than a randomly selected competing source wouldSupported

ChatGPT's ~48% Wikipedia citation share describes topics Wikipedia actually covers well. Nobody has published what wins the citation for the much larger set of topics it doesn't.

Share on X

Why this matters for niche topics

A brand or publisher in a specialized or emerging space can't realistically compete with Wikipedia's citation dominance on general-knowledge queries. But most of their actual content sits in the gap where Wikipedia has nothing.

That's precisely where this study aims to identify what kind of content wins. It's a far more actionable target than trying to "beat Wikipedia."

What to do about it right now

You don't need this study's final result to act sensibly today. If your topic has no real Wikipedia article, aim to become the comprehensive reference page Wikipedia would have written if it covered your space. Structure, definitions, history, and key facts, laid out clearly, the same way an encyclopedia entry would.

Avoid a narrow, single-purpose product page competing for this slot. A broad, well-organized reference page has a much better shot at winning a gap citation than a page that only pitches one product.

Limitations

  • "Thin Wikipedia coverage" requires a defined threshold, which will be specified and published before data collection to prevent post-hoc redefinition.
  • Results are specific to ChatGPT — the same question for Perplexity's Reddit dependency is registered as a separate, parallel study.
How to cite this
Namdev, R. (2026). The ChatGPT Wikipedia Dependency (v1). Retrieved from https://ritiknamdev.com/blog/wikipedia-dependency-chatgpt

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Builds directly on the source-skew finding in ChatGPT citation statistics and most-cited domains in AI search.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

FAQ

Frequently asked questions

Why focus on ChatGPT specifically for this question?
Because it shows the strongest reported Wikipedia skew of the major engines, roughly 47.9% of top citations in vendor-reported data. That makes the "what happens when the default isn't available" question most consequential there.
Isn't it obvious that something else gets cited when Wikipedia has nothing?
Obviously something fills the gap. The interesting question is what kind of source. Does it resemble Wikipedia's structural properties, broad, neutral, well-organized, or does it look completely different? That pattern would tell content creators what to emulate for topics an encyclopedia doesn't cover well.
How common are topics without Wikipedia coverage?
Very common, outside general-knowledge domains. Most product categories, niche technical topics, and emerging concepts have thin or no Wikipedia coverage. That's exactly why this question has real practical value for a large share of content.
Would a company be better off just trying to get a Wikipedia page created?
For many commercial topics, no. Wikipedia's notability guidelines and neutral-point-of-view policy make it a poor fit for product pages, brand pages, and most commercial content. That's a large part of why this gap exists structurally, not by oversight. The more actionable path for most brands is winning the gap-fill citation this study is designed to characterize, not competing to become the encyclopedic default itself.
Could this study's findings vary a lot by topic category?
Plausibly. A gap in coverage of an emerging technology concept, and a gap in coverage of a niche consumer product, likely get filled by different kinds of sources: technical documentation and forums, versus retailer and review sites. Breaking results down by topic category, not just reporting one blended finding, is part of the planned analysis.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.