Reported data puts Wikipedia at nearly 48% of ChatGPT's top citations. That statistic describes topics where Wikipedia exists and is comprehensive. For the large share of topics where it doesn't — niche products, emerging concepts, specific technical questions — what replaces it? This page registers a study to find out, and sits inside the same programme as the AI citation index and the parallel Reddit-dependency study on Perplexity.
No data has been collected yet. Nothing on this page is a result. The gap-topic query set has not been built and no citations have been collected.
- ChatGPT's widely-quoted ~48% Wikipedia citation share describes topics Wikipedia covers well. Its provenance is traced in full on this page.
- Wikipedia's structural properties — broad coverage, neutral tone, stable URLs — are observable and shared by few other sources.
- What wins the citation on the much larger set of topics Wikipedia does not cover.
- Whether citation concentration is lower on gap topics than on covered ones.
- Whether the skew reflects a hard-coded preference or the properties Wikipedia happens to have.
The question nobody has answered
Every published Wikipedia-skew statistic implicitly describes a subset of queries: ones where Wikipedia has relevant, substantial coverage. It says nothing about the queries where that isn't true. Outside general encyclopedic knowledge, that's most of them.
Say a content creator's topic genuinely has no strong Wikipedia article. Knowing what wins the citation in that gap is more useful to them than knowing Wikipedia's overall dominance.
Scale is what makes it matter. ChatGPT's monthly user numbers, tracked by Statista and cross-checked against Similarweb, put it well ahead of the other answer engines. A source-selection quirk in a small tool is trivia; the same quirk in the most-used answer engine is infrastructure.
What is known
ChatGPT's citations reportedly skew toward Wikipedia and encyclopedic sources at roughly 47.9% of top citations — Partial , vendor-reported from a non-public corpus.
What specifically wins the citation for topics without meaningful Wikipedia coverage — no published study addresses this directly.
of ChatGPT's top citations reportedly go to Wikipedia and encyclopedic sources. Vendor-reported from a corpus nobody outside the vendor can inspect.
Independent write-ups point the same way without settling the number. Ziptie on how ChatGPT chooses its sources and Discovered Labs on citation patterns across ChatGPT, Claude and Perplexity both describe an encyclopedic pull on general-knowledge queries. Neither publishes a corpus, which is why every figure on this site carries a verification grade beside it.
Why Wikipedia wins when it's available
Four properties make Wikipedia the default. Broad, reasonably neutral coverage. A predictable structure of definition, history, key facts, references. Frequent updates. And sheer scale — an article already exists for an enormous share of general-knowledge queries.
Any source filling the gap would plausibly need some subset of those properties: structure and reasonable comprehensiveness, if not scale. That is why the hypotheses below predict official documentation and established publications, rather than arbitrary sources, as the likely runners-up.
The prediction sits close to what the original GEO paper argued about how generative systems select sources, and to Onely's account of what makes content LLM-friendly.
Defining "thin coverage" precisely
To avoid a post-hoc, cherry-picked definition of "gap topic," this study specifies a concrete threshold before data collection begins. A topic qualifies as thin-coverage if no Wikipedia article exists for it directly. It also qualifies if the closest matching article mentions the topic only in passing. A single sentence or list entry within a much broader article doesn't count as a dedicated subject.
This threshold gets published up front. Describing it only after seeing which topics produced interesting results wouldn't keep the finding honest. Publishing it first is what keeps this a genuine test, not a selectively illustrated one.
The actual research question
Restrict the analysis to queries about topics with thin or no Wikipedia coverage. Which domain types, structural properties, and content formats capture the citation instead? Does the "runner-up" source resemble Wikipedia's properties, neutral tone, broad structure, comprehensive coverage, or does it look like something else entirely?
| Property | Wikipedia-covered topic | Gap topic |
|---|---|---|
| Default source | A dedicated encyclopedia article exists and is usually current | None; the slot is contested by whatever else covers the subject |
| Likely competitor set | Effectively one dominant source plus supporting citations | Documentation, trade publications, retailers, forums, in unknown proportion |
| Neutrality of the winner | Editorially enforced, non-commercial by policy | Frequently commercial, with no neutrality requirement at all |
| Realistic strategy | Compete on adjacent questions rather than the definitional one | Build the reference page an encyclopedia would have written |
| Evidence available | Vendor-reported citation-share figures, ungraded corpora | None published; this is the hole the study is registered to fill |
Study design
- 01 Identify topics Where Wikipedia has no dedicated article
- 02 Query ChatGPT Across the fixed query set for those topics
- 03 Log citations What replaces the encyclopedic default
- 04 Classify sources By domain type and structure
- 05 Compare Against topics where Wikipedia does exist
- v0Design stage
Threshold, query set and hypotheses published before collection
Registering the definition first is what keeps this a test rather than an illustration.
- v1Collection
Gap topics and matched Wikipedia-covered controls queried on a fixed schedule
Repeated sampling, because single runs are unstable.
- v2Classification
Cited domains coded by type and structural properties
Two coders, with disagreement rates reported.
- v3Publication
Results, nulls and raw files released together
A contrary result is published with the same prominence as a confirmation.
Two design choices matter most, because they are where studies like this usually go wrong. The query set is fixed before collection and repeated on a schedule — single-run answers are unstable, and a one-shot sample finds whatever the sampler hoped for. The classification scheme for cited domains is written down in advance, the same discipline the whole Index programme works under.
Pre-registered hypotheses
| # | Hypothesis | Predicted outcome |
|---|---|---|
| WD1 | In the absence of Wikipedia coverage, official documentation and established industry publications capture a disproportionate share of citations | Expect to hold |
| WD2 | Citation concentration is lower (more distinct domains cited) for gap-topic queries than for Wikipedia-covered ones | Expect to hold |
| WD3 | The runner-up source's structural properties (broad coverage, neutral tone) more closely resemble Wikipedia than a randomly selected competing source would | Expect to hold |
ChatGPT's ~48% Wikipedia citation share describes topics Wikipedia actually covers well. Nobody has published what wins the citation for the much larger set of topics it doesn't.
Share on XOne more prediction is worth recording, because it is cheap to check. Suppose gap citations go to whatever most resembles an encyclopedia entry. The winning pages should then look like the ones described in how to get cited by ChatGPT — direct definitions, self-contained facts, clean headings — rather than simply the pages that rank well. Ahrefs' analysis of which pages AI Overviews cite found the overlap with top-ranking pages weaker than expected, which is consistent with structure mattering independently of position.
Why an encyclopedia is structurally easy to cite
The skew is often described as a preference. It is more useful to read it as five ordinary properties stacking up, none of which requires deliberate favouring.
| Property | Why it favours citation | Where alternatives fail |
|---|---|---|
| Coverage breadth | A source that exists for nearly every topic wins by availability alone | Most reference sites cover one domain — which is exactly the gap this study targets |
| Predictable structure | Definition, then history, then detail is close to ideal for a system extracting a short factual claim; it does not have to hunt | Commercial pages bury the definition below a pitch |
| Licence and accessibility | Openly licensed, unpaywalled, technically simple to fetch | Many high-quality alternatives fail on at least one of the three |
| Neutral register | Prose that sells nothing is safer to quote than a commercial page making the same claim | No neutrality requirement applies to the sources contesting a gap topic |
| Link density | Among the most linked-to domains on the web, so any retrieval system using authority signals surfaces it | A new reference page starts with none of it |
Link density deserves a caveat. Wikipedia is among the most linked-to domains on the web, so any retrieval system using authority signals surfaces it. Whether links or unlinked prominence carry more of that weight is genuinely contested. The correlational case is set out in brand mentions versus backlinks, and all of that evidence is observational.
Hypothesis We expect these properties, not a hard-coded preference, explain most of the skew. One caveat: "this property looks like it should matter" has a poor track record here. The schema markup study is registered precisely because that reasoning failed for llms.txt. The testable consequence is that gap-topic citations should go to whichever source best approximates the same stack — hypothesis WD3.
Where the 47.9% figure actually comes from
A number gets stated with its provenance or not at all, so: the headline figure is vendor-reported and the corpus behind it is not public. We cannot check the query set, the sampling dates, the country, or how a "top citation" was defined. That is why it carries a Partial grade rather than a stronger one.
We use it anyway, deliberately: the direction is corroborated by ordinary observation — ask a general-knowledge question and encyclopedic sources appear constantly. It is the decimal we do not trust, not the pattern.
One structural detail deserves more attention than the decimal. ChatGPT's web results have historically leaned on Bing's index, and Seer Interactive found that 87% of SearchGPT citations matched Bing's top results. If that relationship holds, part of the encyclopedic skew is inherited from a conventional search index rather than chosen by the model. That is a different claim with different implications.
None of the gap-topic design depends on the figure being 47.9 rather than 40 or 55. The study measures its own baseline from its own query set, so a badly wrong vendor figure leaves it standing. That independence was built in deliberately.
The circularity problem nobody wants to discuss
One structural issue underneath all of this deserves stating plainly, even though this study cannot resolve it. Wikipedia articles are built from citations to other sources. If AI answers cite Wikipedia heavily, and people increasingly get their facts from AI answers, then a large share of public knowledge flows through one editorial pipeline. That pipeline has known coverage biases: by language, by geography, by topic, and by who volunteers to edit.
A second-order version is worse. If AI-generated text spreads across the web, and Wikipedia editors later cite web sources that were themselves AI-generated from Wikipedia, the loop closes. A claim could eventually be supported only by its own echo.
Open question We have no measurement of how much of this is happening, and we are not going to speculate with numbers. It sits next to another unmeasured question: how long a cited page stays cited, the subject of the citation half-life study. It matters most where being wrong is expensive, which is why YMYL content in AI search is tracked separately. The mechanism is not exotic and the incentives point toward it growing — and no citation-share statistic captures any of it.
Confounds in a gap-topic study
Suppose the study finds that documentation and established publications fill the gap. Before believing WD1, several alternative explanations have to be ruled out.
| Confound | Why it could produce the same result |
|---|---|
| Topic-age imbalance | Gap topics skew newer. Newer topics have different source ecosystems regardless of Wikipedia's absence. |
| Commercial density | Many gap topics are commercial, where documentation and retailer content simply dominate the available web. |
| Query phrasing | Gap-topic queries may be phrased more specifically, which favours documentation over general reference by lexical match alone. |
| Wikipedia's own quality gradient | Topics with no article often also have thinner coverage everywhere. Scarcity, not substitution, may drive the result. |
| Selection of the gap set | If the thin-coverage threshold is applied loosely, the set fills with topics chosen for interestingness rather than by rule. |
The first two are the serious ones. Both would produce exactly the predicted pattern with no substitution behaviour involved at all. Matching gap and non-gap topics on age and commercial intent is therefore not optional. It is what makes the comparison mean anything.
What does not transfer to other engines
Everything on this page is about one engine, deliberately. Different engines draw on different source pools and different retrieval architectures. A skew toward community forums on one engine and toward encyclopedic sources on another are not two views of one phenomenon; they are different systems making different choices. Per-engine citation shares differ enough that a blended "AI search" figure hides more than it shows. Cross-platform citation concordance is the protocol registered to measure how small the overlap really is, and Ahrefs reached a similar conclusion measuring overlap between AI search engines.
Gap behaviour is even less transferable than the baseline: what fills a hole depends on what is nearby in that engine's source pool. An engine leaning on community content fills a gap with community content; one leaning on a classic web index fills it with whatever that index ranks.
Within Google alone the point holds. AI Mode and AI Overviews are not one system, and the source delta between them has its own registered protocol.
Hypothesis We expect gap-fill behaviour to vary more between engines than baseline behaviour does, because gaps expose the underlying source pool rather than the shared default. That is a prediction the parallel study on a second engine is designed to test.
How to check your own topic for a Wikipedia gap
You can do a rough version of this analysis for your own topic set in an hour. It will not be publishable research, but it will tell you which strategy applies to you.
- List your twenty most commercially important topics as plain-language subjects, not keywords.
- For each, check whether a dedicated Wikipedia article exists, applying the same threshold this study uses: a passing mention inside a broader article does not count.
- Sort into three buckets — strong article, thin or passing mention, nothing at all.
- For the gap buckets, ask an AI engine the natural question a reader would ask and record which sources it names. Repeat on a different day; single samples are unreliable.
- Look at what the cited sources have in common structurally, not just who they are — broad reference pages, documentation, comparison articles. That shape is your target.
Before any of it, confirm the engine can reach you at all. Most AI crawlers do not execute JavaScript, so a reference page whose text arrives client-side is invisible however encyclopedic it reads. Check that retrieval bots are allowed through in robots.txt, then run the technical GEO audit over the page itself.
Two categories need separate treatment. Local queries resolve against a different source pool entirely, and shopping queries pull retailer and review content no encyclopedia was going to supply. Treat the whole exercise as reconnaissance — no control group, no repeated sampling, no statistical power — which is still more than most content planning starts with.
Null results we would publish
Four outcomes would be null or contrary, and each gets published with the same prominence as a confirmation, in the null results registry:
- Gap-topic citations show no coherent structural pattern at all — which would mean the advice to build encyclopedic reference pages has no support from this line of evidence, worth knowing before anyone spends a year on it.
- Citation concentration is the same or higher for gap topics, contradicting WD2.
- The runner-up source resembles Wikipedia no more than a random competitor does, contradicting WD3.
- A definitional failure: a thin-coverage threshold two people cannot apply consistently to the same topic list. Boring, and genuinely useful — it would tell everyone else attempting this study where the work actually is.
Open questions
Open question How much of a gap-fill citation is explained by page structure versus domain authority? The two are almost always confounded in observation.
Open question Does the encyclopedic skew differ by language, and how badly does it hurt topics whose best sources are not in English?
Open question When Wikipedia coverage for a topic improves, does citation shift toward it, and how quickly? A natural experiment nobody has run.
Open question How much of the web that engines cite is already downstream of Wikipedia? The circularity question, restated as something measurable in principle.
This study is one slice of the AI Citation Index programme. The crawl-to-citation latency study answers the timing half of the gap question: how long after publishing a reference page you could expect to see anything.
No gap-topic queries have been run. The results, the topic list and the raw citation records go out to the newsletter when the first collection window closes. To propose a gap topic for the set, or to point out a flaw in the design before it runs, how to reach me is on the about page — there is no sign-up form, just a conversation.
Limitations
- "Thin Wikipedia coverage" requires a defined threshold, which will be specified and published before data collection to prevent post-hoc redefinition.
- Results are specific to ChatGPT — the same question for Perplexity's Reddit dependency is registered as a separate, parallel study.
- The baseline figure is vendor-reported, from a corpus nobody outside the vendor can inspect, and is used here as a direction rather than a measurement.
- Gap and non-gap topics must be matched on age and commercial intent, or the two most serious confounds above reproduce the predicted result with no substitution behaviour involved.
- English-language, single-region collection. Wikipedia's coverage depth varies enormously by language, so a gap in the English encyclopedia is not a gap everywhere, and no figure here should be read as describing non-English retrieval.
- Topic selection is a judgement call. The thin-coverage threshold is published in advance, but which twenty verticals the gap topics are drawn from will shape the runner-up mix. The topic list ships with the data so the choice can be argued with.
- Measurement is of what the answer surface displays, not of what the model retrieved. A source consulted but not named leaves no trace, so citation share is a floor on source use, not a full account of it.
- Reproduction is approximate. Anyone can re-run the published query set, but a later run measures a later product state and a later encyclopedia; agreement on direction, not on the decimal, is the realistic bar.
Namdev, R. (2026). The ChatGPT Wikipedia Dependency (v1). Retrieved from https://ritiknamdev.com/blog/wikipedia-dependency-chatgpt Published under CC BY 4.0 — reuse freely with attribution.
Builds directly on the source-skew finding in ChatGPT citation statistics and most-cited domains in AI search.