Original research · Pre-registered

The Query Fan-Out Corpus

Everyone repeats that Google AI Mode fans a query into '8-16 sub-queries.' Nobody has published the data behind that number. This is the design for a corpus that would.

The single highest-originality asset registered in the Citation Index roadmap. Replaces industry folklore with an inferred, published, checkable dataset — with its inference limitations stated up front, not hidden.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — design stage ·15 min read ·Last verified September 2026
The short version

Google documents that AI Mode "may use a query fan-out technique," and a granted patent describes eight sub-query type categories. The specific number of sub-queries issued per real question — the "8-16" figure that circulates constantly — has never been published with a disclosed method. This corpus is designed to infer and publish that data directly, with its limitations stated plainly.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. The corpus has not been collected. Sub-query counts below are inferred design targets, not measurements.

What is known
  • Google documents that AI Mode decomposes a query into sub-queries.
  • The widely-repeated 8–16 sub-query range could not be traced to any disclosed measurement — that trace is on this page.
What is not yet known
  • How many sub-queries are actually inferred for real questions.
  • Whether sub-query count scales with question complexity.
  • Whether pages covering multiple angles appear across more sub-query clusters.

What question does this corpus answer?

How many sub-queries does Google AI Mode actually issue for a real question, and what determines the number? Nobody knows. "AI Mode fans a query out into 8–16 sub-queries" appears across dozens of AI SEO articles, always with the same confidence and never with a visible source. Traced back, the underlying patent describes eight sub-query types, not a count of instances per query — and the repeated figure appears to have mixed those two things up.

This page registers a corpus designed to replace that folklore with an inferred, published, checkable dataset. It does not assume the folklore is wrong; it replaces the assumption with a measurement, as a slice of the AI Citation Index. Nothing has been collected.

Anatomy of a number with no source

The "8 to 16" figure is a useful specimen, because it shows exactly how a number acquires authority without ever acquiring evidence.

It begins with something real: a patent names eight categories of sub-query. That is a documented fact about a document. Then a category count gets read as an instance count — eight types becomes eight queries, and nothing in the source supports the swap. Then a range appears. "Eight to sixteen" sounds more measured than a flat eight, and ranges feel like they came from data; this one did not, as far as we can trace. Then repetition does the rest: each restatement drops a qualifier, and within a few hops the number is presented as an established property of the system.

StepWhat is actually trueWhat gets said
The sourceA patent names eight categories of sub-query"Google says eight sub-queries"
The swapA taxonomy of typesA count of searches issued
The rangeNo measurement exists"Eight to sixteen", which sounds derived from data
The repetitionEach restatement drops a qualifierAn established property of the system
The attributionIndustry shorthandAttributed to Google directly

The lesson generalises: when a figure sounds specific but never carries a method, look for the moment a category became a count. That is the ordinary life cycle of an unsourced number in this field, documented at length in where AI SEO statistics come from.

What is actually documented?

Fact

Google's own documentation confirms the fan-out mechanism exists, and a granted patent (US11663201B2, filed 2018) names eight sub-query type categories. Google's separate AI optimization guidance is the nearest thing to official advice on any of it.

Open question

The actual number of sub-queries issued per real-world question, and how it varies with question complexity, is not published anywhere we could locate with a disclosed method.

What fan-out is, mechanically

Strip away the jargon and it is a simple idea: one question from a person becomes several questions asked internally, and the answers are combined into one response. "Is this camera worth upgrading to?" hides at least four questions — what is it, how does it compare, what does it cost, what else do I need. A single retrieval pass answers that badly; splitting it retrieves better material for each part.

The consequence for a publisher is that you are not competing for one slot but for a set of narrower slots, most of which you never see. That is a materially different game from the one measured by click-through and zero-click statistics, and part of why AI referral traffic is so hard to attribute to a cause.

Open question How the split is decided, how many parts are produced, and whether the parts are fixed or generated per question, are not publicly documented for the shipping product.

Method: how a fan-out will be captured

Inference pipeline
  1. 01 Issue query From the published 1,000-query set
  2. 02 Capture full response Including any observable sub-answer structure
  3. 03 Infer sub-query boundaries From citation clusters and answer segments
  4. 04 Classify each sub-query Against the 8 documented patent categories
  5. 05 Publish the corpus Query → inferred sub-queries → citations, open

AI Mode does not expose an internal sub-query list through any known public interface. The corpus therefore works from observable proxies: distinct answer segments, non-overlapping citation clusters, and structural cues in the response such as headings, comparison tables and clear topic shifts. Each segment is treated as one inferred sub-query and classified against the patent's eight documented types. These are inferred boundaries, not a direct readout of Google's internal process, and every record in the published corpus is labelled inferred rather than observed.

Sample: how the query set gets built

A published 1,000-query set, constructed to stated rules before the hypotheses are tested, and released with the corpus. Composition matters more than size. A set full of comparison questions will find comparison sub-queries, so a neutral set has to include the boring cases: simple factual questions where decomposition would be pointless, plus local and commercial questions, where behaviour is known to differ. Each query is issued multiple times in clean sessions from a fixed region inside a dated collection window, so non-determinism is measured rather than absorbed.

Variables: what is observable, and what is not

Answer segmentsDistinct topic blocks in the response, marked by hand or by rule.
Citations per segmentWhich sources a segment draws on, and which it shares with its neighbours.
Repeat-run variationHow much of the structure survives asking the same question again.
The internal sub-query listNever exposed through any public interface we know of.
Sub-queries that returned nothingThey leave no trace, so inference undercounts them systematically.
Whether structure reflects retrievalA response template would look identical to a retrieval trace.

Recorded per query: the query text, run number, timestamp, the full response, each inferred segment with its boundary evidence, the ordered citation list per segment, and the patent-category classification for each segment. Two of the unobservable items above bias the result in a known direction — sub-queries that returned nothing leave no trace, so inference undercounts systematically, and a response template would look identical to a retrieval trace. Both are reported as bounds, not corrected away.

Pre-registered hypotheses

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
FO1Inferred sub-query count correlates positively with question complexity (measured by word count and entity count)Expect to hold
FO2The commonly-cited "8-16" range is roughly consistent with observed inferred countsExpect to hold
FO3Pages covering multiple angles of a topic (comparison + specification + pricing) appear in more inferred sub-query clusters than single-angle pagesExpect to hold

FO2 exists to give the folklore a fair test rather than assume it wrong by default. If the repeated range turns out to be roughly accurate, that is worth publishing plainly; replacement is not the goal, an accurate answer is. Comparable work on the AI Overviews side — Semrush's AI Overviews study and Ahrefs on how often citations come from the top 10 — measures nothing about fan-out, but shows what a disclosed method looks like. The difference between the two Google surfaces is the subject of the AI Mode versus AI Overviews source delta.

'AI Mode fans a query into 8-16 sub-queries' is repeated everywhere with no visible source. We're not assuming it's wrong — we're building the corpus that would let anyone actually check.

Share on X

Control and validation: how the inference gets checked

Blind human coding. A subset of responses is segmented independently by a person blind to the automated inference, and the two are compared. The agreement rate is published whether or not it is flattering; low agreement means the automated method is revised before the corpus is trusted. Blind double-coding on a sample is also the only real defence against coder drift, since a person segmenting hundreds of answers will change their standard over time.

Repeat-run stability. The same query runs multiple times to test whether the inferred structure stays consistent. Wildly different structures for the same question would indicate the method is picking up noise rather than an underlying pattern, and that floor is reported alongside every count.

A negative control. The query set deliberately includes simple factual questions where decomposition should be minimal. If the method infers as many segments there as for genuinely multi-part questions, it is measuring writing style, not retrieval. Both checks are ordinary research hygiene, the same ones applied in the schema markup study and the content freshness study. A method that cannot report its own error rate should not be believed, including ours.

Why counting sub-queries is harder than it sounds

Even with perfect access, counting would be contested. Without it, several judgement calls have to be made explicit in advance.

What counts as one sub-query? A retrieval call, a reasoning step, or a distinct topic in the answer — these give different counts for the same response. Do repeated calls count twice? A system may issue the same sub-query more than once, or refine it, and counting attempts is not the same measure as counting distinct questions. Are unanswered sub-queries visible? One that returned nothing may leave no trace in the output at all. Does the answer preserve the structure? A response written as flowing prose can merge several sub-answers, making segment boundaries genuinely ambiguous.

For that reason the corpus reports inferred segment counts and states clearly that they are a lower bound on whatever the system does internally. Two further traps sit alongside. Treating format as structure: three headings may reflect a writing style rather than three retrievals. And over-trusting citation clusters: two segments can share sources because the sources are broad, not because the segments are one sub-query.

Confounds in the inference

Several things could produce a segment pattern that has nothing to do with fan-out. Answer length limits — a truncated answer shows fewer segments regardless of how many sub-queries ran. Presentation templates — if the product formats certain question types into fixed sections, the structure is a template, not a retrieval trace. Model verbosity changes — an update that makes answers shorter would look exactly like a drop in sub-query count. Personalisation — different accounts or locations may get different structure for the same question.

Topic availability is the fifth. On a thin topic there may be nothing to retrieve for some sub-queries, collapsing the answer — the subject of the thin-topic study on another engine.

Surface drift is the confound a long collection window cannot escape: the share of searches running through this product is itself moving, as market-share statistics show. A corpus collected over a long window is measuring a moving object, which is why the window is kept short and dated, and a meaningful product change is logged as a distinct era rather than blended with earlier data.

How should a fan-out count be read?

Four misreadings are predictable enough to rule out now. "Fan-out means I need sixteen pages" — no count implies a page count, and the relationship between sub-queries and pages is precisely what is unmeasured. "The patent describes the live product" — a granted patent describes an invention that may or may not ship in the form described. "More segments means a better answer" — segment count is a structural observation and says nothing about quality. "If I cover all eight categories I will be cited" — a hypothesis with no supporting evidence, and one of the things this corpus would let someone test.

Nor does a count transfer between engines. Other systems decompose questions too, but the number of parts may differ and the categories almost certainly do, because the patent categories reflect one company's product thinking. How far apart the engines are is itself partly measured — Ahrefs on overlap between AI search engines finds substantial divergence, and the concordance study is our own version of the question. Treat any fan-out figure as engine-specific and dated.

Hypothesis The one thing worth acting on before the corpus exists is not an optimisation but a writing discipline: list the follow-up questions a reader would obviously ask, and check whether anything you publish answers them. It has a real cost — bloat, dilution, and pricing sections that go stale — and it is defensible on its own merits rather than on any sub-query count. Restructuring a site around an unverified number is the risk this whole page exists to name. It also assumes the basics: a page an AI crawler cannot render is not a candidate for any sub-query, which is why a technical GEO audit comes first.

Topics that naturally decomposeProducts, services, technical decisions with a chain of follow-ups.
One-page-versus-cluster decisionsExactly where fan-out evidence would change an editorial choice.
Simple factual topicsOne right answer, little to decompose.
Audiences off the surfaceCheck your readers use the product before optimising for its mechanism.

Null results we would publish

  • No relationship between complexity and segment count. FO1 predicts one. If it is absent, the prediction fails publicly.
  • Counts far outside the folklore range. If inferred counts sit well below or above eight to sixteen, we publish the discrepancy rather than softening it.
  • An unreliable inference method. If human coders disagree with the automated segmentation, we publish the disagreement rate and hold the corpus back.
  • The eight patent categories do not fit real answers. A patent filed years ago describes an idea, not necessarily the shipping system. We report the misfit rather than forcing observations into categories that do not match.
  • No advantage to multi-angle pages. FO3 is the most commercially interesting prediction. A null there would contradict a great deal of current advice, including advice we find plausible.

Each would be filed in the null results registry alongside every other prediction on this site that did not survive contact with data.

Three things would change this design outright. Official documentation of how many sub-queries the live product issues. An interface exposing the decomposition, which would turn inference into observation and make this design obsolete in the best possible way. Or evidence that the answer's visible structure is generated after retrieval and independently of it, which would invalidate segment-based inference entirely.

Schedule

Study phases, as registered
  1. Design publishedNow

    This page

    Hypotheses, query set rules and classification scheme fixed before any data exists.

  2. Collection and inferencePlanned

    Dated collection window

    Repeated runs per query, so non-determinism is measured rather than absorbed.

  3. ValidationPlanned

    Blind human sample

    Agreement rate published whether or not it is flattering.

  4. Corpus releasePlanned

    Open dataset plus method

    Query, inferred segments, citations per segment, downloadable and disputable.

Design first, data second, conclusion last — the same sequencing as the citation dataset strategy, with every study on the roadmap listed under studies and the reporting conventions taken from the AI visibility measurement standard.

Raw data and reproduction instructions

The published artefact is the full corpus — query → inferred sub-query segments → citations per segment — as a downloadable file. It is released with the inference methodology and the query set, so the classification logic can be audited and individual records disputed. A published corpus with a stated method can be corrected by someone else; a number with no method can only be repeated.

You cannot see the internal queries, but you can learn a great deal from the output cheaply:

Observe fan-out behaviour on one question yourself
  1. 1 Pick one question with obvious parts Something involving a comparison, a price and a requirement works well.
  2. 2 Ask it three times in clean sessions Save the full answers, including every source named. One run is an observation, not a measurement.
  3. 3 Mark the topic shifts by hand Where the answer stops discussing one thing and starts another is your segment boundary.
  4. 4 List the sources per segment Note which segments share sources and which do not. Non-overlap is the strongest available signal of a separate retrieval.
  5. 5 Ask the parts separately Run each sub-question as its own query and compare its sources to the combined answer. If your page appears standalone but not combined, that gap is the interesting thing.
  6. 6 Record the date and exact wording Without both you cannot repeat this next quarter and compare. Dated repetition is the whole point.

Open questions

  • Does the same question asked in two languages decompose the same way?
  • Are sub-queries generated fresh per question, or drawn from a fixed set of templates by question type?
  • How often does a sub-query return nothing, and does the system tell the user?
  • Does a page cited for one sub-query gain any advantage for the others?
  • Is there a ceiling on decomposition, and does hitting it change answer quality?

Limitations

  • Inference, not observation. Without access to AI Mode's internal query log, sub-query boundaries are reconstructed from output. Every count is a lower bound containing some misclassification, and the corpus labels each record inferred.
  • Engine coverage. One surface of one company's product. Nothing here describes how ChatGPT, Perplexity, Claude or Gemini decompose a question, and the patent categories are unlikely to fit them.
  • Query selection. The 1,000-query set drives the result. A set weighted toward comparison questions would find more segments; the published set lets anyone test how much of the headline depends on that choice.
  • Geography and language. English-language queries from a fixed region and a clean session. Whether decomposition differs by language or locale is unmeasured, and listed as an open question rather than answered.
  • Measurement. Segment boundaries are contested even among human coders; the blind-coding agreement rate quantifies that disagreement rather than removing it.
  • Confounders. Response templates, verbosity changes and answer-length limits can all produce segment patterns unrelated to retrieval. The design detects none of these directly.
  • Reproducibility. The product can change behaviour without notice, so every count is a dated snapshot rather than an architectural fact, and each collection window is published as its own era.
  • Generalisation. A segment count describes output structure. It licenses no claim about how many pages to write, about citation odds, or about answer quality.
Where to go next

To see how many distinct sub-questions your own topic gets decomposed into: run the output-side procedure above on ten of your own queries and count the answer segments. To read the collection programme this corpus feeds: the sampling frame is in the AI Citation Index.

When this runs

The corpus, the query set and the inference code go out to the newsletter when the first collection window closes. If you want to argue with the segmentation rules or propose queries for the set before it is locked — the point at which the objection is most useful — how to reach me is on the about page.

How to cite this
Namdev, R. (2026). The Query Fan-Out Corpus (v1). Retrieved from https://ritiknamdev.com/blog/query-fan-out-corpus-study

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

See Google AI Mode statistics and query fan-out, explained for the mechanism this corpus will measure directly, and the AI Citation Index for where this study sits in the wider roadmap.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Search Engine Journal — query fan-out technique in AI Modewww.searchenginejournal.com/query-fan-out-technique-in-ai-mode-new-details-from-google/552532 US Patent 11663201B2patents.google.com/patent/US11663201B2 Google Search Central — AI features and your websitedevelopers.google.com/search/docs/appearance/ai-features Google Search Central — AI optimization guidedevelopers.google.com/search/docs/fundamentals/ai-optimization-guide Aggarwal et al. — GEO: Generative Engine Optimization (arXiv)arxiv.org/abs/2311.09735 Aggarwal et al. — GEO paper, full PDFarxiv.org/pdf/2311.09735 Wikipedia — Generative engine optimizationen.wikipedia.org/wiki/Generative_engine_optimization Semrush — AI Overviews studywww.semrush.com/blog/semrush-ai-overviews-study Ahrefs — how often AI Overview citations come from the top 10ahrefs.com/blog/ai-overview-citations-top-10 Search Engine Journal — AI Overview citations from top-ranking pages drop sharplywww.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637 Ahrefs — overlap between AI search enginesahrefs.com/blog/ai-search-overlap Ahrefs — AI SEO statisticsahrefs.com/blog/ai-seo-statistics Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors Ziptie — how original research wins AI citationsziptie.dev/blog/how-original-research-wins-ai-citations Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Discovered Labs — how ChatGPT, Claude and Perplexity choose sourcesdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Salespeak — content freshness and AI searchsalespeak.ai/aeo-news/content-freshness-ai-search Similarweb — generative AI usage statisticsaisearch.similarweb.com/blog/gen-ai-stats StatCounter — search engine market sharegs.statcounter.com/search-engine-market-share Ritik Namdev — where AI SEO statistics come fromritiknamdev.com/blog/where-ai-seo-statistics-come-from
FAQ

Frequently asked questions

Can you directly observe Google's internal sub-queries?
Not directly. AI Mode doesn't expose its internal sub-query list through any public interface we're aware of. This corpus infers sub-query structure from observable signals, answer segmentation, citation clustering, response structure, rather than from privileged internal access. That is a real methodological limitation, stated plainly in the limitations section.
Why does the "8-16 sub-queries" figure need replacing at all?
Because it's repeated across dozens of SEO articles with no visible source, and our own trace couldn't find a primary study behind it. See the mechanism explainer for the full trace. It may be roughly accurate, or it may not be. Nobody has published the data to check.
What would this corpus let someone do that they can't do today?
Test, for the first time with real data instead of assumption, whether covering a topic's adjacent sub-angles, comparisons, specifications, pricing, measurably increases the chance of being pulled into a fan-out's sub-query results.
How will you know if the inference method is actually reliable?
Through the validation approach registered above. It cross-checks inferred sub-query boundaries against independent human review of a sample, and tests whether the same query produces consistent inferred structure across repeated runs.
Could this corpus become outdated quickly if Google changes AI Mode?
Yes, and that's treated as an expected part of the research program, not a flaw. Each collection window is dated. A meaningful product change gets logged as a distinct era in the corpus, rather than silently blended with earlier data.
Does a higher sub-query count mean more chances to be cited?
That is the intuitive reading, and it may be wrong. More sub-queries also means each one is narrower, so a general page may fit none of them well. Whether breadth or depth wins is exactly the kind of question the corpus is meant to make testable rather than arguable.
Should I write one long page covering every angle, or several focused pages?
Nobody can answer that from public evidence today. Both structures have a plausible mechanism behind them. The honest position is that this is an open question, and anyone giving you a confident answer is reasoning from intuition rather than data.
Is inferring sub-queries from the output a legitimate method?
It is legitimate if its limits are stated and its reliability is checked, which is why the validation step exists. It is not a substitute for direct observation, and the corpus will describe every record as inferred rather than observed. A method with known error is more useful than no method, provided the error is disclosed.
What happens if the eight patent categories do not fit real answers?
Then the classification scheme is wrong for the current product, and that is a finding. A patent filed years ago describes an idea, not necessarily the shipping system. We would report the misfit rather than forcing observations into categories that do not match.
Could this study reach the wrong number and still be useful?
Yes. A published corpus with a stated method can be corrected by someone else. A number with no method cannot be corrected at all, only repeated. Being checkably wrong is more valuable to the field than being unaccountably confident.
If I cannot see the sub-queries, is there any point acting on this at all?
There is a limited one. You can list the follow-up questions a reader would obviously ask, and check whether anything you publish answers them clearly. That is a writing discipline rather than an optimisation, and it holds its value whatever the eventual corpus shows.
Why publish a study design before running the study?
Because a prediction written afterwards is not a prediction. Publishing the design fixes the hypotheses, the query set and the classification rules before any data exists, so nobody, including us, can quietly adjust them to fit a result that reads better.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.