Google documents that AI Mode "may use a query fan-out technique," and a granted patent describes eight sub-query type categories. The specific number of sub-queries issued per real question — the "8-16" figure that circulates constantly — has never been published with a disclosed method. This corpus is designed to infer and publish that data directly, with its limitations stated plainly.
No data has been collected yet. Nothing on this page is a result. The corpus has not been collected. Sub-query counts below are inferred design targets, not measurements.
- Google documents that AI Mode decomposes a query into sub-queries.
- The widely-repeated 8–16 sub-query range could not be traced to any disclosed measurement — that trace is on this page.
- How many sub-queries are actually inferred for real questions.
- Whether sub-query count scales with question complexity.
- Whether pages covering multiple angles appear across more sub-query clusters.
What question does this corpus answer?
How many sub-queries does Google AI Mode actually issue for a real question, and what determines the number? Nobody knows. "AI Mode fans a query out into 8–16 sub-queries" appears across dozens of AI SEO articles, always with the same confidence and never with a visible source. Traced back, the underlying patent describes eight sub-query types, not a count of instances per query — and the repeated figure appears to have mixed those two things up.
This page registers a corpus designed to replace that folklore with an inferred, published, checkable dataset. It does not assume the folklore is wrong; it replaces the assumption with a measurement, as a slice of the AI Citation Index. Nothing has been collected.
Anatomy of a number with no source
The "8 to 16" figure is a useful specimen, because it shows exactly how a number acquires authority without ever acquiring evidence.
It begins with something real: a patent names eight categories of sub-query. That is a documented fact about a document. Then a category count gets read as an instance count — eight types becomes eight queries, and nothing in the source supports the swap. Then a range appears. "Eight to sixteen" sounds more measured than a flat eight, and ranges feel like they came from data; this one did not, as far as we can trace. Then repetition does the rest: each restatement drops a qualifier, and within a few hops the number is presented as an established property of the system.
| Step | What is actually true | What gets said |
|---|---|---|
| The source | A patent names eight categories of sub-query | "Google says eight sub-queries" |
| The swap | A taxonomy of types | A count of searches issued |
| The range | No measurement exists | "Eight to sixteen", which sounds derived from data |
| The repetition | Each restatement drops a qualifier | An established property of the system |
| The attribution | Industry shorthand | Attributed to Google directly |
The lesson generalises: when a figure sounds specific but never carries a method, look for the moment a category became a count. That is the ordinary life cycle of an unsourced number in this field, documented at length in where AI SEO statistics come from.
What is actually documented?
Google's own documentation confirms the fan-out mechanism exists, and a granted patent (US11663201B2, filed 2018) names eight sub-query type categories. Google's separate AI optimization guidance is the nearest thing to official advice on any of it.
The actual number of sub-queries issued per real-world question, and how it varies with question complexity, is not published anywhere we could locate with a disclosed method.
What fan-out is, mechanically
Strip away the jargon and it is a simple idea: one question from a person becomes several questions asked internally, and the answers are combined into one response. "Is this camera worth upgrading to?" hides at least four questions — what is it, how does it compare, what does it cost, what else do I need. A single retrieval pass answers that badly; splitting it retrieves better material for each part.
The consequence for a publisher is that you are not competing for one slot but for a set of narrower slots, most of which you never see. That is a materially different game from the one measured by click-through and zero-click statistics, and part of why AI referral traffic is so hard to attribute to a cause.
Open question How the split is decided, how many parts are produced, and whether the parts are fixed or generated per question, are not publicly documented for the shipping product.
Method: how a fan-out will be captured
- 01 Issue query From the published 1,000-query set
- 02 Capture full response Including any observable sub-answer structure
- 03 Infer sub-query boundaries From citation clusters and answer segments
- 04 Classify each sub-query Against the 8 documented patent categories
- 05 Publish the corpus Query → inferred sub-queries → citations, open
AI Mode does not expose an internal sub-query list through any known public interface. The corpus therefore works from observable proxies: distinct answer segments, non-overlapping citation clusters, and structural cues in the response such as headings, comparison tables and clear topic shifts. Each segment is treated as one inferred sub-query and classified against the patent's eight documented types. These are inferred boundaries, not a direct readout of Google's internal process, and every record in the published corpus is labelled inferred rather than observed.
Sample: how the query set gets built
A published 1,000-query set, constructed to stated rules before the hypotheses are tested, and released with the corpus. Composition matters more than size. A set full of comparison questions will find comparison sub-queries, so a neutral set has to include the boring cases: simple factual questions where decomposition would be pointless, plus local and commercial questions, where behaviour is known to differ. Each query is issued multiple times in clean sessions from a fixed region inside a dated collection window, so non-determinism is measured rather than absorbed.
Variables: what is observable, and what is not
Recorded per query: the query text, run number, timestamp, the full response, each inferred segment with its boundary evidence, the ordered citation list per segment, and the patent-category classification for each segment. Two of the unobservable items above bias the result in a known direction — sub-queries that returned nothing leave no trace, so inference undercounts systematically, and a response template would look identical to a retrieval trace. Both are reported as bounds, not corrected away.
Pre-registered hypotheses
| # | Hypothesis | Predicted outcome |
|---|---|---|
| FO1 | Inferred sub-query count correlates positively with question complexity (measured by word count and entity count) | Expect to hold |
| FO2 | The commonly-cited "8-16" range is roughly consistent with observed inferred counts | Expect to hold |
| FO3 | Pages covering multiple angles of a topic (comparison + specification + pricing) appear in more inferred sub-query clusters than single-angle pages | Expect to hold |
FO2 exists to give the folklore a fair test rather than assume it wrong by default. If the repeated range turns out to be roughly accurate, that is worth publishing plainly; replacement is not the goal, an accurate answer is. Comparable work on the AI Overviews side — Semrush's AI Overviews study and Ahrefs on how often citations come from the top 10 — measures nothing about fan-out, but shows what a disclosed method looks like. The difference between the two Google surfaces is the subject of the AI Mode versus AI Overviews source delta.
'AI Mode fans a query into 8-16 sub-queries' is repeated everywhere with no visible source. We're not assuming it's wrong — we're building the corpus that would let anyone actually check.
Share on XControl and validation: how the inference gets checked
Blind human coding. A subset of responses is segmented independently by a person blind to the automated inference, and the two are compared. The agreement rate is published whether or not it is flattering; low agreement means the automated method is revised before the corpus is trusted. Blind double-coding on a sample is also the only real defence against coder drift, since a person segmenting hundreds of answers will change their standard over time.
Repeat-run stability. The same query runs multiple times to test whether the inferred structure stays consistent. Wildly different structures for the same question would indicate the method is picking up noise rather than an underlying pattern, and that floor is reported alongside every count.
A negative control. The query set deliberately includes simple factual questions where decomposition should be minimal. If the method infers as many segments there as for genuinely multi-part questions, it is measuring writing style, not retrieval. Both checks are ordinary research hygiene, the same ones applied in the schema markup study and the content freshness study. A method that cannot report its own error rate should not be believed, including ours.
Why counting sub-queries is harder than it sounds
Even with perfect access, counting would be contested. Without it, several judgement calls have to be made explicit in advance.
What counts as one sub-query? A retrieval call, a reasoning step, or a distinct topic in the answer — these give different counts for the same response. Do repeated calls count twice? A system may issue the same sub-query more than once, or refine it, and counting attempts is not the same measure as counting distinct questions. Are unanswered sub-queries visible? One that returned nothing may leave no trace in the output at all. Does the answer preserve the structure? A response written as flowing prose can merge several sub-answers, making segment boundaries genuinely ambiguous.
For that reason the corpus reports inferred segment counts and states clearly that they are a lower bound on whatever the system does internally. Two further traps sit alongside. Treating format as structure: three headings may reflect a writing style rather than three retrievals. And over-trusting citation clusters: two segments can share sources because the sources are broad, not because the segments are one sub-query.
Confounds in the inference
Several things could produce a segment pattern that has nothing to do with fan-out. Answer length limits — a truncated answer shows fewer segments regardless of how many sub-queries ran. Presentation templates — if the product formats certain question types into fixed sections, the structure is a template, not a retrieval trace. Model verbosity changes — an update that makes answers shorter would look exactly like a drop in sub-query count. Personalisation — different accounts or locations may get different structure for the same question.
Topic availability is the fifth. On a thin topic there may be nothing to retrieve for some sub-queries, collapsing the answer — the subject of the thin-topic study on another engine.
Surface drift is the confound a long collection window cannot escape: the share of searches running through this product is itself moving, as market-share statistics show. A corpus collected over a long window is measuring a moving object, which is why the window is kept short and dated, and a meaningful product change is logged as a distinct era rather than blended with earlier data.
How should a fan-out count be read?
Four misreadings are predictable enough to rule out now. "Fan-out means I need sixteen pages" — no count implies a page count, and the relationship between sub-queries and pages is precisely what is unmeasured. "The patent describes the live product" — a granted patent describes an invention that may or may not ship in the form described. "More segments means a better answer" — segment count is a structural observation and says nothing about quality. "If I cover all eight categories I will be cited" — a hypothesis with no supporting evidence, and one of the things this corpus would let someone test.
Nor does a count transfer between engines. Other systems decompose questions too, but the number of parts may differ and the categories almost certainly do, because the patent categories reflect one company's product thinking. How far apart the engines are is itself partly measured — Ahrefs on overlap between AI search engines finds substantial divergence, and the concordance study is our own version of the question. Treat any fan-out figure as engine-specific and dated.
Hypothesis The one thing worth acting on before the corpus exists is not an optimisation but a writing discipline: list the follow-up questions a reader would obviously ask, and check whether anything you publish answers them. It has a real cost — bloat, dilution, and pricing sections that go stale — and it is defensible on its own merits rather than on any sub-query count. Restructuring a site around an unverified number is the risk this whole page exists to name. It also assumes the basics: a page an AI crawler cannot render is not a candidate for any sub-query, which is why a technical GEO audit comes first.
Null results we would publish
- No relationship between complexity and segment count. FO1 predicts one. If it is absent, the prediction fails publicly.
- Counts far outside the folklore range. If inferred counts sit well below or above eight to sixteen, we publish the discrepancy rather than softening it.
- An unreliable inference method. If human coders disagree with the automated segmentation, we publish the disagreement rate and hold the corpus back.
- The eight patent categories do not fit real answers. A patent filed years ago describes an idea, not necessarily the shipping system. We report the misfit rather than forcing observations into categories that do not match.
- No advantage to multi-angle pages. FO3 is the most commercially interesting prediction. A null there would contradict a great deal of current advice, including advice we find plausible.
Each would be filed in the null results registry alongside every other prediction on this site that did not survive contact with data.
Three things would change this design outright. Official documentation of how many sub-queries the live product issues. An interface exposing the decomposition, which would turn inference into observation and make this design obsolete in the best possible way. Or evidence that the answer's visible structure is generated after retrieval and independently of it, which would invalidate segment-based inference entirely.
Schedule
- Design publishedNow
This page
Hypotheses, query set rules and classification scheme fixed before any data exists.
- Query set constructionNext
1,000 queries
Built to the stated rules, including questions where fan-out would be pointless.
- Collection and inferencePlanned
Dated collection window
Repeated runs per query, so non-determinism is measured rather than absorbed.
- ValidationPlanned
Blind human sample
Agreement rate published whether or not it is flattering.
- Corpus releasePlanned
Open dataset plus method
Query, inferred segments, citations per segment, downloadable and disputable.
Design first, data second, conclusion last — the same sequencing as the citation dataset strategy, with every study on the roadmap listed under studies and the reporting conventions taken from the AI visibility measurement standard.
Raw data and reproduction instructions
The published artefact is the full corpus — query → inferred sub-query segments → citations per segment — as a downloadable file. It is released with the inference methodology and the query set, so the classification logic can be audited and individual records disputed. A published corpus with a stated method can be corrected by someone else; a number with no method can only be repeated.
You cannot see the internal queries, but you can learn a great deal from the output cheaply:
- 1 Pick one question with obvious parts Something involving a comparison, a price and a requirement works well.
- 2 Ask it three times in clean sessions Save the full answers, including every source named. One run is an observation, not a measurement.
- 3 Mark the topic shifts by hand Where the answer stops discussing one thing and starts another is your segment boundary.
- 4 List the sources per segment Note which segments share sources and which do not. Non-overlap is the strongest available signal of a separate retrieval.
- 5 Ask the parts separately Run each sub-question as its own query and compare its sources to the combined answer. If your page appears standalone but not combined, that gap is the interesting thing.
- 6 Record the date and exact wording Without both you cannot repeat this next quarter and compare. Dated repetition is the whole point.
Open questions
- Does the same question asked in two languages decompose the same way?
- Are sub-queries generated fresh per question, or drawn from a fixed set of templates by question type?
- How often does a sub-query return nothing, and does the system tell the user?
- Does a page cited for one sub-query gain any advantage for the others?
- Is there a ceiling on decomposition, and does hitting it change answer quality?
Limitations
- Inference, not observation. Without access to AI Mode's internal query log, sub-query boundaries are reconstructed from output. Every count is a lower bound containing some misclassification, and the corpus labels each record inferred.
- Engine coverage. One surface of one company's product. Nothing here describes how ChatGPT, Perplexity, Claude or Gemini decompose a question, and the patent categories are unlikely to fit them.
- Query selection. The 1,000-query set drives the result. A set weighted toward comparison questions would find more segments; the published set lets anyone test how much of the headline depends on that choice.
- Geography and language. English-language queries from a fixed region and a clean session. Whether decomposition differs by language or locale is unmeasured, and listed as an open question rather than answered.
- Measurement. Segment boundaries are contested even among human coders; the blind-coding agreement rate quantifies that disagreement rather than removing it.
- Confounders. Response templates, verbosity changes and answer-length limits can all produce segment patterns unrelated to retrieval. The design detects none of these directly.
- Reproducibility. The product can change behaviour without notice, so every count is a dated snapshot rather than an architectural fact, and each collection window is published as its own era.
- Generalisation. A segment count describes output structure. It licenses no claim about how many pages to write, about citation odds, or about answer quality.
To see how many distinct sub-questions your own topic gets decomposed into: run the output-side procedure above on ten of your own queries and count the answer segments. To read the collection programme this corpus feeds: the sampling frame is in the AI Citation Index.
The corpus, the query set and the inference code go out to the newsletter when the first collection window closes. If you want to argue with the segmentation rules or propose queries for the set before it is locked — the point at which the objection is most useful — how to reach me is on the about page.
Namdev, R. (2026). The Query Fan-Out Corpus (v1). Retrieved from https://ritiknamdev.com/blog/query-fan-out-corpus-study Published under CC BY 4.0 — reuse freely with attribution.
See Google AI Mode statistics and query fan-out, explained for the mechanism this corpus will measure directly, and the AI Citation Index for where this study sits in the wider roadmap.