Statistics · ChatGPT

ChatGPT citation statistics

ChatGPT's distinctive source signature — an encyclopedic skew and a Bing-derived index lineage — with the population behind each figure attached, and the one crawler switch that decides whether you are eligible at all.

Ritik Namdev Ritik Namdev ·Published September 2026 ·Platform-specific ·12 min read ·Last verified September 2026
The short version

What is distinctive about ChatGPT is its encyclopedic source signature sitting on a Bing-derived index. Around 47.9% of its top citations reportedly go to Wikipedia and similarly encyclopedic sources. That is the highest single-source concentration reported for any major engine, and it describes a query mix weighted toward topics an encyclopedia covers well. Google ranking barely predicts it: about 12% overlap with Google's top 10, against ~38% for Google's own AI Overviews. And the switch that decides whether you are eligible at all is OAI-SearchBot, not GPTBot.

What this page establishes
  • ~47.9% of ChatGPT top citations reportedly go to Wikipedia and similarly encyclopedic sources — vendor corpus of a stated 680M citations, query mix undisclosed, reported 2026. Graded partial. The caveat travels with it: that share describes topics an encyclopedia covers well.
  • ~12% of URLs cited by ChatGPT also rank in Google's top 10 for the same prompt — Ahrefs, 2026, grouped with Gemini and Copilot. Against ~33% for Perplexity and ~38% for AI Overviews. Graded traceable.
  • One analysis found 87% of SearchGPT citations matching Bing's top results (Seer Interactive, method described, data not published), which makes Bing visibility a load-bearing input most teams never check.
  • OAI-SearchBot performs the live fetch that produces a citation; GPTBot crawls for training. Blocking the second protects nothing about today's visibility. Blocking the first removes you from ChatGPT Search entirely.
  • No published figure states how often a ChatGPT query triggers retrieval at all — the denominator under every citation rate on this page. Two studies with different query mixes will disagree for reasons unrelated to the engine.

What does ChatGPT actually cite?

47.9%

of ChatGPT's top citations reportedly go to Wikipedia and similarly encyclopedic sources — the largest single-source concentration reported for any major engine.

Profound — graded partial · Reported 2026
~12%

of URLs cited by ChatGPT (grouped with Gemini and Copilot) also rank in Google's top 10 for the same prompt.

Ahrefs — graded traceable · 2026
Reported split of ChatGPT's top citations
  • Wikipedia and similarly encyclopedic sources 47.9
  • Everything else — the tail every other publisher competes in 52.1
Vendor-reported, from a stated 680-million citation corpus that is not independently public. The ordering is more reliable than the decimal place.

Independent teardowns of how ChatGPT selects sources, how it differs from Claude and Perplexity and the platform-level citation patterns all describe the same shape without publishing a corpus anyone can recompute. The most-cited domains page holds the per-engine comparison, and the concordance study runs a fixed panel instead of a vendor corpus.

The Wikipedia skew, and the caveat that travels with it

Wikipedia's prominence in ChatGPT citations is one of the most reliably reproduced observations in this field, and one of the most frequently misused. The figure is meaningless without the caveat, so they should never be separated: a 47.9% encyclopedic share is measured across a query mix weighted toward topics an encyclopedia covers well. On broad definitional questions an encyclopedia is close to an ideal retrieval source. On specific, current, practical or data-bearing questions it has nothing to offer at all, and the share on those questions is not 47.9% — it is unmeasured, because nobody has published the query-type breakdown.

Hypothesis Why the skew exists: Wikipedia is broad, structured, neutral in register, densely interlinked, freely licensed and unusually consistent in format. Almost every property you would design into an ideal retrieval source, it has. That is inference from the observed pattern, not a mechanism OpenAI has described.

What it does not mean is that you should try to be Wikipedia, or that editing Wikipedia is a visibility tactic — that is against its rules and it backfires. Read as a map of where not to fight, the skew is genuinely useful. The remaining share is the tail where every non-encyclopedic source competes, and it is reached by being specific and current rather than comprehensive and neutral. The Wikipedia dependency study asks the sharper version of the question — what ChatGPT cites when the encyclopedia has nothing on a topic — and the gaps it finds are where an ordinary publisher can win.

~47.9% of ChatGPT's top citations reportedly go to encyclopedic sources — on a query mix weighted toward topics an encyclopedia covers well. Quote the caveat or don't quote the number.

Share on X

The Bing lineage: ChatGPT's most under-checked input

ChatGPT Search performs live retrieval at query time rather than answering purely from training data, and it has historically drawn in part on a Bing-derived index alongside OpenAI's own crawl infrastructure. The exact current architecture is not fully disclosed and has likely evolved. The lineage is not speculation, though. Seer Interactive found 87% of SearchGPT citations matching Bing's top results, which makes Bing's index and user base far more relevant to ChatGPT visibility than its consumer market share suggests. The Bing and Copilot guide covers that surface on its own terms.

This is the distinctive fact about ChatGPT among the assistants, and the practical one. A team optimising hard for Google while never checking Bing Webmaster Tools is neglecting the input with the tightest reported relationship to its citations. Retrieval also depends on the page being fetchable at all — whether these systems execute JavaScript is covered in the rendering study, and the hygiene checklist is the technical GEO audit.

Why Google ranking barely predicts a ChatGPT citation

At roughly 12%, Google ranking is only weakly associated with a ChatGPT citation — a sharp contrast with AI Overviews at about 38%, itself down sharply year over year. Four reasons make that plausible rather than surprising.

Different index. ChatGPT's retrieval has not historically run on Google's, and two corpora with different crawl priorities disagree about what exists before they disagree about what is good. Different unit. Google ranks pages against a query; retrieval selects passages that support composing an answer, and a page can be an excellent supporting passage and a mediocre standalone result.

Fan-out. Retrieval is often for a sub-question rather than the typed query, so comparing citations against rankings for the original query compares two different questions. The fan-out corpus study measures that directly. And no click competition: Google's ranking has been shaped for years by what people click, and retrieval has no equivalent pressure.

The practical conclusion is uncomfortable and clear: your Google rankings are a weak proxy for your AI visibility, and using one to report the other will mislead you in both directions.

GPTBot vs OAI-SearchBot: the switch that matters

GPTBotOAI-SearchBot
What it doesCrawls to improve future model trainingFetches live, at answer time
Effect on today's citationsNoneThis is the switch
Reason to blockConsent and compensation for your workRarely a good one for a public page
Cost of blockingZero current visibilityRemoval from ChatGPT Search citations
Most common mistakeBlocking it and believing you protected visibilityBlocking it by accident, via an old wildcard rule

Two crawlers means two separate decisions, and conflating them is the most common error we see. A site that blocks OAI-SearchBot while allowing GPTBot removes itself from ChatGPT Search citations while still contributing to future model training — likely the opposite of its owner's intent.

Both tokens are documented by OpenAI in the platform bots page and the developer docs; the consolidated list across every vendor is the bot user-agent registry. How many sites pull each lever is measured in the robots.txt blocking census and independently in publisher-blocking studies; how much traffic each represents is in the crawler statistics hub and Cloudflare's crawler breakdown.

Four stages of ChatGPT web access, and what each one changed
  1. Stage 1Early browsing

    A limited, optional web-access feature

    Citation patterns documented under this stage do not describe the current product. Old screenshots circulate anyway.

  2. Stage 2GPTBot published

    A named training crawler with a documented robots.txt token

    This is the crawler most site owners blocked. It has no bearing on whether you are cited today.

  3. Stage 3ChatGPT Search

    A dedicated search product with its own retrieval and citation UI

    OAI-SearchBot performs the live fetch that actually produces a citation.

That lineage matters when reading older statistics or screenshots. A citation pattern documented under the early browsing feature does not describe current ChatGPT Search behaviour, and the figures on this page describe the current-generation product as of their stated dates.

When ChatGPT does not search at all

Not every answer involves retrieval. The model can answer from training, and when it does there is no source list and no citation opportunity for anyone. That matters more than it first appears. A citation rate computed over a query set including many non-searching answers measures two things at once, and two studies with different query mixes will disagree for reasons unrelated to either engine or either publisher.

It also bounds what optimisation can achieve. On questions answered from training, no on-page work reaches the user, because the page is never fetched. Leverage concentrates on questions specific, current or contested enough to trigger a search. When you build your own panel, mark which answers searched — it turns a noisy average into two clean ones.

Cited, linked, named: three different outcomes

Almost every citation-rate statistic collapses three outcomes worth very different amounts. Retrieved means your page was fetched and considered — invisible to the user, invisible to you, and producing nothing on its own. Linked means your page appears in the source list; a user who looks at sources might click, and most do not. Named means your brand is mentioned inside the answer text, which is the outcome with real value because it reaches the reader who never opens a source list.

Almost no published study separates linked from named, and named is the one a marketing team actually wants. When you run your own measurement, record them in separate columns. It costs nothing and it is the single biggest improvement you can make over the public figures.

Referral volume, meanwhile, stays a small fraction of total traffic for most sites even where conversion quality is reported higher. The conversion benchmarks page and the referral traffic page handle those claims properly, since the published multiples vary too widely to restate as one number.

Measuring your own ChatGPT presence

A repeatable five-step ChatGPT visibility panel
  1. 1 Pick 20-50 real questions Customer wording from tickets, sales calls and query reports. Not internal wording.
  2. 2 Run each three times Answers vary between runs. One run tells you nothing.
  3. 3 Record five fields Did it search; were you linked; were you named; who else was cited; what did it get wrong.
  4. 4 Hold conditions constant Same account state, same region, both written down.
  5. 5 Repeat quarterly, keep raw records A snapshot is a curiosity. Four quarters is evidence.

Twenty to fifty real questions in customer wording, from support tickets and sales calls rather than internal phrasing. Each run three times, because variance between runs is real and one run tells you nothing. Five fields recorded per run: did it search, were you linked, were you named, who else was cited, and what did it get wrong. Conditions held constant and written down — same account state, same region. Repeated quarterly with the raw records kept, because the value is entirely in the trend. A single snapshot is a curiosity; four quarters is evidence, and the reporting conventions are in the measurement standard.

When you report upward, lead with the trend rather than the level. Show the competitor column: who else got cited on your questions is far more persuasive than your own count. And name what you are not claiming — attributed revenue, above all. Propose a review date rather than a decision, since this channel does not yet support confident budget calls.

What your server logs can and cannot tell you

Because OpenAI runs separately named crawlers, log analysis is more informative here than for Google's surfaces — the structural reason Gemini research is so much thinner. It still has hard limits.

Logs can tell you that a fetch happened, which agent made it, which URL it wanted and what status you returned. That is enough to catch the two most common real problems: blocking something you did not mean to block, and returning errors to a crawler while serving humans fine. Logs cannot tell you whether the fetch produced a citation, whether you were named, or how many people read the answer. There is no path from a request line to an outcome.

And verify user agents properly. A user-agent string is self-reported and trivially forged. Confirm against published address ranges before treating a log line as evidence, and be sceptical of any crawler-traffic figure that skipped this step.

Five mistakes in reading these numbers

Treating a source-mix study as a target list. Knowing which domains dominate tells you where the competition is, not where to publish.

Comparing figures across studies. Different query panels, dates, regions and counting rules produce numbers that are not comparable, however similar they look.

Reading correlation as instruction. Cited pages sharing a trait does not mean the trait caused the citation — the standing problem across the tactic evidence scoreboard.

Assuming stability. These products change without announcement. Our working rule: source-mix figures stay directionally useful for roughly six to twelve months, and specific percentages should be treated as historical after that. Anything about crawler names or tokens should be checked against primary documentation before you rely on it, because that layer changes fastest and breaks most silently.

Ignoring the denominator. A citation rate over queries where the engine never searched is measuring something else entirely.

Open questions

Open question How often does a query trigger retrieval? The denominator under most citation statistics, and nobody has published it.

Open question What is the naming rate versus the linking rate? Measurable with a modest panel. Not published anywhere we have found.

Open question How stable are citations for the same query over time? The citation half-life study is the attempt to measure it, and the latency study asks it from the other end. Our first-party zero-to-cited log study is where both started.

Open question Does interface matter? Web, mobile and API access may retrieve differently. Registered as open, because we have no evidence either way.

Next step

Check what your robots directives actually expose — the single most common ChatGPT visibility failure is an old wildcard rule blocking OAI-SearchBot, and it costs nothing to rule out. New per-engine measurements ship through the newsletter.

Verification status

The ranking-overlap figure traces to Ahrefs with a disclosed sample — Traceable . The source-skew figure is vendor-reported from a non-public corpus — Partial — and should never be quoted without the query-mix caveat above. Conversion claims are treated as Broken chain per the provenance audit until a reconciled, like-for-like figure exists. The same scheme runs on GEO statistics and AEO statistics; the standards behind it are in the method notes.

How to cite this
Namdev, R. (2026). ChatGPT citation statistics (v2). Retrieved from https://ritiknamdev.com/blog/chatgpt-citation-statistics

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

This page is a slice of the AI Citation Index, the recurring measurement it reports into. For the crawler mechanics behind the data, see GPTBot vs OAI-SearchBot; for the tactical playbook, how to get cited by ChatGPT.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Ahrefs — only 12% of AI-cited URLs rank in Google's top 10ahrefs.com/blog/ai-search-overlap Ahrefs — AI Overview citations and top-10 rankingsahrefs.com/blog/ai-overview-citations-top-10 Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns OpenAI — GPTBot and OAI-SearchBot documentationplatform.openai.com/docs/bots OpenAI developer docs — bots and user agentsdevelopers.openai.com/api/docs/bots Similarweb — AI search traffic and referral dataaisearch.similarweb.com/blog/gen-ai-stats Statista — Global monthly ChatGPT userswww.statista.com/statistics/1659718/global-monthly-chatgpt-users Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Backlinko — Bing user statisticsbacklinko.com/bing-users Ziptie — How does ChatGPT choose its sourcesziptie.dev/blog/how-does-chatgpt-choose-its-sources Discovered Labs — How ChatGPT, Claude and Perplexity choose sourcesdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Discovered Labs — How each platform cites sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Leapd — How ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Search Engine Journal — Query fan-out in AI Mode, new details from Googlewww.searchenginejournal.com/query-fan-out-technique-in-ai-mode-new-details-from-google/552532 Aggarwal et al. (arXiv) — GEO: Generative Engine Optimizationarxiv.org/abs/2311.09735 Cloudflare — From Googlebot to GPTBot: who is crawling your siteblog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Cloudflare — Crawl-to-click ratios for AI botsblog.cloudflare.com/crawlers-click-ai-bots-training Search Engine Journal — Cloudflare report, Googlebot tops AI crawler trafficwww.searchenginejournal.com/cloudflare-report-googlebot-tops-ai-crawler-traffic/563303 Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots Paul Calvano — AI bots and robots.txtpaulcalvano.com/2025-08-21-ai-bots-and-robots-txt BuzzStream — Publishers blocking AI crawlers studywww.buzzstream.com/blog/publishers-block-ai-study
FAQ

Frequently asked questions

Does ChatGPT use Google or Bing as its index?
ChatGPT Search has historically leaned on a Bing-derived index for parts of its retrieval, alongside OpenAI's own crawling via OAI-SearchBot. The exact current split is not publicly disclosed and likely shifts over time. Treat any specific claim about which index ChatGPT uses as approximate.
Why is Wikipedia cited so heavily?
Plausibly because ChatGPT's retrieval leans toward broad, structured, neutral-register reference content, and Wikipedia is close to an ideal fit. It also reflects the query mix behind the measurement: an encyclopedia dominates on topics an encyclopedia covers well, and has nothing to offer on specific, current or practical questions. This is inference from the observed pattern, not a mechanism OpenAI has confirmed.
Does ranking well on Google help get cited by ChatGPT?
Only weakly, per available data. Roughly 12% of URLs cited by ChatGPT, grouped with Gemini and Copilot in the cited study, also ranked in Google's top 10 for the same prompt. Google ranking is a much stronger predictor for AI Overviews specifically than for ChatGPT.
What is the difference between GPTBot and OAI-SearchBot for citation purposes?
GPTBot crawls to improve future model training and has no direct bearing on today's citations. OAI-SearchBot performs the live retrieval that actually produces a citation in ChatGPT Search. Blocking OAI-SearchBot removes you from ChatGPT Search citations. Blocking GPTBot does not.
If I only do one thing for ChatGPT visibility, what should it be?
Confirm OAI-SearchBot is not blocked. Everything else is optimisation; that one is a gate. Teams discover a stray robots.txt rule from years ago far more often than you would expect.
Why do I get cited for a question one day and not the next?
Generated answers vary between runs even with identical inputs, and the retrieved source set varies with them. A single absence is not evidence of anything. Only a sustained pattern across repeated runs and many queries is a signal.
Should I write content aimed specifically at ChatGPT?
No. Write content aimed at being the clearest, best-sourced answer to a real question. That is what the available evidence points at, and unlike a platform-specific format it does not become worthless when the platform changes.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.