Reported data puts Reddit at nearly 47% of Perplexity's top citations. But that describes topics with substantial community discussion. For B2B, enterprise, and specialized topics where Reddit discussion is thin, what wins the citation instead? This page registers a study to find out. It mirrors the design used for ChatGPT's Wikipedia dependency.
No data has been collected yet. Nothing on this page is a result. Query set and collection window are specified below. Nothing has been measured yet.
- Reddit appears prominently in reported Perplexity source mixes. The figures are vendor-reported and graded Partial here.
- Perplexity's source mix visibly differs from ChatGPT's on the same questions.
- How large the dependency actually is on a defined query set.
- Whether it holds across query categories or concentrates in a few.
- Whether community presence is therefore a usable visibility channel — currently a Hypothesis, not a finding.
What question does this study answer?
When a topic has no substantial Reddit discussion, what does Perplexity cite instead? Nobody has published an answer. The reported Reddit skew describes topics where community discussion is dense — consumer products, hobbies, mainstream software. A large share of B2B, enterprise and regulated professional content generates almost none of it, and for those topics the entire published picture of how this engine sources answers simply does not apply.
That gap is where most commercial visibility questions actually live, which is why it is worth a protocol rather than an opinion. This page registers the design before collection, as a slice of the AI Citation Index and a direct counterpart to the Wikipedia dependency study: identify a gap where an engine's dominant source has nothing, then measure what fills it. Running both under one design also makes a later comparison possible — whether the two engines resolve an absent default the same way is a question neither study can answer alone.
Perplexity's citations reportedly skew toward Reddit at roughly 46.7% of top citations — Partial , vendor-reported from a non-public corpus.
What wins the citation for topics with thin Reddit discussion — no published study addresses this.
Where does the ~47% Reddit figure come from?
of Perplexity's top citations reportedly go to Reddit — the figure this entire study exists to qualify rather than repeat.
Traced back, it originates with Profound's citation-pattern corpus, and is then restated in Discovered Labs' platform comparison and Leapd's sourcing summary. The corpus is not public, the collection date is not stated, and the query panel behind it is not described — so it is graded Partial by the measurement standard, and the tracing method is the provenance audit.
Independent write-ups of how each engine selects sources — Discovered Labs, Ziptie and Zyppy's correlational study of citation factors — converge on the same description without re-measuring it. That is corroboration, not replication. The wider figures for this engine sit in Perplexity citation statistics, and the cross-engine version of the table is the most-cited-domains page.
How is "thin discussion" defined?
This is the study's weakest joint: if the definition is fuzzy, everything downstream inherits the fuzziness. So the threshold is fixed before collection and applied without exceptions.
Count discussions, not mentions. A topic mentioned in passing across many threads is not the same as a topic with its own threads. The second is what a retrieval system can use. Count substantive replies. A question posted with no answer is not discussion; a minimum number of substantive responses is required before a thread counts.
Verify a sample by hand. Automated counts misclassify, and reading a subsample to report the error rate is the only honest way to know how badly. Publish the topic list. Anyone should be able to disagree with our classification of a specific topic and check what that does to the result.
A threshold adjusted after seeing results is not a threshold. Where the thin regions tend to fall is predictable enough to map in advance, which is also how the topic pool gets stratified:
Method: how the citations get collected
- 01 Identify topics Where Reddit has thin or no discussion
- 02 Query Perplexity Across the fixed query set for those topics
- 03 Log citations What replaces the community-discussion default
- 04 Classify sources By domain type and content structure
- 05 Compare Against topics where Reddit discussion is dense
Every query is run from clean sessions with no account history, in a fixed region, inside a single collection window, and the full source list is logged rather than the top source only. Collection also depends on being able to see what the engine fetches, which is less settled than it sounds: Perplexity documents its own crawlers, while Cloudflare has reported undeclared crawling attributed to it. The access side is covered by the bot registry and the robots.txt blocking census.
Sample and control: thin topics against a dense baseline
The sample is a fixed query set over topics classified thin by the threshold above, stratified across the categories in the map — enterprise procurement, regulated practice, narrow technical fields, very new topics. The query set is published with the results, not selected after seeing them.
The control is a matched set of Reddit-dense topics run in the same window under identical conditions. Without it, a low Reddit share on thin topics is uninterpretable: it could reflect the thinness, or the panel, or the day. A second control runs a subset of queries twice, because if answers vary between runs then any single-run source mix is partly noise, and the size of that noise floor has to be reported next to every difference the study claims.
Variables: what gets recorded per query
Pre-registered hypotheses
| # | Hypothesis | Predicted outcome |
|---|---|---|
| RD1 | In the absence of Reddit discussion, first-party vendor and documentation content captures a disproportionate share of citations on Perplexity | Expect to hold |
| RD2 | B2B/enterprise queries show meaningfully lower Reddit citation share than consumer queries in the baseline (Reddit-present) comparison group | Expect to hold |
RD1 is a reasoned bet about what is typically available for thin B2B topics, not a foregone conclusion. Review-site dominance and analyst-report dominance are both plausible alternative outcomes the data could support instead, and the analysis below treats them as equal candidates.
Perplexity's ~47% Reddit citation share describes topics with active community discussion. Most B2B and enterprise content doesn't have that — and nobody has published what wins the citation there instead.
Share on XAnalysis: four substitution outcomes, four different conclusions
The core comparison is the domain-type distribution on thin topics against the same distribution on the dense control, tested as a difference in shares with a confidence interval and read against the repeat-run noise floor. Four outcomes are possible, and each leads somewhere different:
Another community platform would mean the preference is for the community format rather than for one site, and little changes for publishers. Vendor and specialist pages would be the commercially significant outcome, because thin regions would then be winnable by ordinary publishers.
Weak, distant threads — the engine stretching to use a barely relevant thread rather than switch source type — would be a structural preference strong enough to hurt answer quality. That is a finding about the engine, not about publishing. Nothing useful at all, where the engine hedges or declines, is informative too, and quite likely for the narrowest questions. Each outcome is reported as a share, not narrated as a story.
Three mechanisms behind the skew
Why community discussion dominates is much less settled than that it does, and the cause determines whether anything can be done about it.
Availability. Threads exist for an enormous range of questions, phrased the way people actually ask them. For many long-tail questions there is nothing else on the open web addressing them directly. Format fit. A thread is a question followed by several answers with visible agreement signals, which maps unusually well onto what a retrieval system needs — compare a marketing page, where the answer is buried under positioning.
Access arrangements. Some platforms hold commercial agreements with AI companies, and those agreements affect what a system can retrieve and reuse. That is a commercial fact, not a quality judgement, and advice that treats source mix as a meritocracy rarely mentions it.
Hypothesis The relative weight of the three is unknown. Our expectation is that availability dominates — which is exactly why the thin-topic case is the informative test: it removes availability while leaving format fit and access arrangements intact. A fourth possibility, brand familiarity independent of any single document, is a separate question from this one.
The behaviour deserves a steelman rather than a diagnosis of defect, because a preference doing real work will not simply be tuned away. A thread opens with a question phrased the way a person would phrase it — a far more direct match to a user query than a page written around a keyword. A single thread can hold agreement, disagreement and edge cases at once, which is efficient for composing a balanced answer in a way a single-voice article is not.
Votes and replies also supply a rough quality estimate that requires no independent judgement about the author, and most open-web writing about a product is trying to sell it while community discussion is comparatively less so. Taken together that is a coherent case. It also means the weaknesses catalogued below come attached to the strengths rather than separable from them.
What a community answer is good and bad at
Criticism of source concentration often slides into dismissing community writing altogether, which is unfair and weakens the argument. The real concern is narrower: a heavy lean toward any single format imports that format's specific weaknesses into every answer built on it.
| Property | Community thread | Vendor documentation |
|---|---|---|
| Lived experience | Strong — the reason it gets cited | Weak, and structurally so |
| Currency | Weak — old answers outrank correct ones | Strong, if maintained |
| Verifiable authorship | Absent — nobody checks credentials | Present, and attributable |
| Numbers with provenance | Unreliable — restated without source | The one thing it can uniquely supply |
| Commercial neutrality | Comparatively high | Low by construction |
| Availability on thin topics | Often none at all | Whatever the vendor chose to publish |
Currency is the row that connects to a measurable question rather than a rhetorical one. Whether recency is rewarded at all is the subject of the freshness study, and how long a citation survives is the half-life study.
How should any source-mix number be read?
This study will produce percentages, and so does everyone else in the area. Four questions decide whether any of them means anything, ours included. What was the query panel? A consumer-heavy panel and a professional panel produce completely different source mixes from the same engine; the panel drives the result more than the engine does. Domain-level or citation-level? The two diverge sharply when one site is cited many times through many pages.
Was variance measured? Without a repeat-run floor, a single-run source mix is partly noise reported as structure. When was it collected? Licensing and access arrangements change, so a source mix from eighteen months ago may describe an arrangement that no longer exists.
A figure that cannot answer all four is a talking point, not a measurement, which is why this design is published before the data. If RD1 holds, the practical reading is narrow. On a thin topic you compete with few pages, possibly none, and the page most likely to win supplies what a thread would have: specifics, real numbers with provenance, honest limits. That is a strategy claim, graded Hypothesis , and it is not evidence until the study runs. What is already known about winning citations at all is in how to get cited and the zero-to-cited log study.
One tactic will not be recommended whatever the data says: seeding community threads about your own product. It violates platform rules, it is increasingly detectable, the reputational cost lands on the brand permanently, and manufactured threads attract little engagement — which makes them the ones least likely to be retrieved in the first place.
The null results we would publish
Written down in advance so they cannot be quietly dropped if the data disappoints. If community sources dominate thin topics just as heavily as dense ones, the substitution hypothesis is wrong and we publish that. If our thinness classification does not correlate with anything in the citations, that is a failure of the method and it goes on the page. If the differences are smaller than run-to-run variance, we report the noise floor and decline to publish a headline figure. If the result merely reproduces existing source-mix work, we say so rather than presenting a replication as a discovery.
All four go into the null-results registry regardless of how uninteresting they look, because a field where only positive findings get published is one where nobody can separate a real effect from a selection artefact.
How stable would any finding here be?
Short shelf life, and four forces would shorten it. Licensing changes start or end, moving the source mix sharply for reasons unrelated to content quality. Platform access policy — a community platform restricting automated access — changes what is retrievable overnight. Quality pressure: if leaning on community sources produces visibly wrong answers, engines have an incentive to rebalance, possibly already and invisibly from outside. And thin topics get less thin as discussion accumulates, which makes any opportunity here real but time-limited.
A fifth force sits outside the engine entirely: the retrieval channel itself may move. Agentic browsers fetch differently from search crawlers, and some agent-facing routes produce no citation at all. Any source-mix finding assumes the current plumbing. For that reason the useful output is not a percentage but a repeated measurement, so direction of travel is visible rather than inferred.
Raw data and reproduction instructions
The published artefact is the per-query CSV — query, topic, thinness classification and its evidence, run number, and the ordered source list — alongside the topic list and the summary. Anyone disagreeing with a thinness call can reclassify from the same file.
You do not need the full study to learn whether the pattern holds in your niche, and the narrow version is the one that affects your decisions:
- 01 List fifteen real questions From tickets and sales calls, in the customer's own words
- 02 Search the community platform Note whether substantive discussion exists — your thinness column
- 03 Run each through the engine Record the full source list, not just the top one
- 04 Compare the two columns Thin questions returning specialist sources means the pattern holds
- 05 Repeat five of them Without a second run you cannot tell a pattern from a coincidence
If your thin questions return specialist sources while your dense ones return threads, the pattern holds in your niche. Run at least five queries twice: answers vary between runs, and without a repeat you cannot tell a pattern from a coincidence.
Limitations
- Sample. A fixed panel of thin-topic queries plus a matched dense control. It is not a random sample of questions asked of this engine, and the panel composition drives the source mix more than the engine does.
- Engine coverage. One engine only. Perplexity is chosen because its source lists are explicit enough to measure cheaply; nothing here should be assumed to describe how another engine handles an absent default.
- Query selection. Thinness is a threshold we set. A different threshold, or a differently stratified topic pool, could move the headline share materially — which is why the topic list ships with the data.
- Geography and language. English-language queries from a fixed region. Non-English markets are probably the largest thin region of all and are the least studied here, by us and by everyone else.
- Measurement. Automated thinness classification will misclassify some topics; the hand-verified subsample gives an error rate, not a correction.
- Confounders. Thin topics differ from dense ones in more ways than thinness — commercial intent, recency, publisher density. The design cannot separate those, so a difference in source mix is not evidence that thinness caused it.
- Reproducibility. Generated answers vary between runs and the engine changes without notice, so any figure is valid for its collection window only and carries its timestamp.
- Baseline. The ~47% skew this study works against is vendor-reported and graded Partial . The study measures the gap; it does not verify the headline.
To test the pattern in your own niche before the study runs: use the seven-step afternoon version above, then check which of your pages are already surfaced with the AI Overview Exposure Checker. To see how the same design behaves on the other engine: read the Wikipedia dependency study for ChatGPT.
The results and the raw per-query files go to the newsletter when the collection window closes, and land alongside the rest of the study programme. If you want to propose a thin topic for the panel, or argue with the thinness threshold before it is locked, how to reach me is on the about page.
Namdev, R. (2026). The Perplexity Reddit Dependency (v0). Retrieved from https://ritiknamdev.com/blog/reddit-dependency-perplexity Published under CC BY 4.0 — reuse freely with attribution.
The Perplexity counterpart to the Wikipedia dependency study. See Perplexity citation statistics for the underlying skew figure.