The question: do Google's two AI surfaces cite the same sources for the same query? Both are documented to use "query fan-out," but they have different interfaces and likely different retrieval depth, and no direct same-query comparison has been published by anyone. This page registers the query set, the four metrics and the predictions before a single paired query has been run.
No data has been collected yet. Nothing on this page is a result. This is the within-Google concordance test. No paired queries have been run.
- AI Mode and AI Overviews are separate products, and coverage routinely conflates them.
- Both surfaces reportedly use query fan-out, so some shared retrieval behaviour is plausible.
- Whether the two surfaces cite the same sources for the same query, and how far they diverge.
- Whether optimising for one surface transfers to the other.
- Whether the delta is stable or moves as both products change.
What question is this study trying to answer?
Take an identical query set, run it through both surfaces under comparable conditions: what share of cited domains overlap, and where they diverge, what characterises the difference?
It is unanswered because nobody has published a direct, same-query comparison. Both surfaces belong to Google and both reportedly use fan-out — a mechanism Google described in some detail for AI Mode specifically. So most coverage treats them as interchangeable, differing only in placement: embedded snippet versus dedicated mode. That conflation is one of the recurring errors documented in the statistics-provenance audit. It has to be got right before any Google figure means anything.
The comparison also only makes sense if both surfaces can reach the page at all, which puts crawler access and robots.txt policy upstream of everything here.
What is known about each surface?
AI Overview citation-to-ranking overlap fell from ~76% to ~38% year over year, per Ahrefs' top-10 citation analysis, reported by SEJ as a sharp drop and documented in the dedicated statistics page.
AI Mode's comparable overlap figure, and how it compares directly to AI Overviews' on the same query set, hasn't been published.
AI Overview citations that also ranked in Google's top 10, year over year. The surface moved away from conventional rankings fast.
AI Mode's comparable overlap figure, on the same query set, in the same window. No public comparison exists.
Why would two Google surfaces cite differently?
| AI Overviews | AI Mode | |
|---|---|---|
| Where it lives | A box above conventional results | A dedicated conversational surface |
| How a user reaches it | Appears, unprompted, on some queries | Entered deliberately |
| Fallback in view | The classic listings, right below | None — the answer is the page |
| Retrieval depth (reported) | Narrow, fast pass | Extensive fan-out across sub-queries |
| Follow-up turns | Not part of the surface | Each turn is another retrieval |
| Publisher control | One coarse AI opt-out, not split by surface | |
Fan-out means a single question becomes several searches. The system breaks the question into narrower parts, searches for each, and assembles one answer from everything it gathered. "Is a standing desk worth it for lower back pain?" might become separate searches for health effects, back pain causes, sit-stand alternation evidence, and typical costs — four searches, four source pools.
A surface that runs closer to one search and summarises the top results draws from a much narrower pool, one that overlaps heavily with what already ranks. That is the crux. If one surface fans out and the other does not, they are not two skins on one system but two retrieval processes that share a company. It is also why Semrush's AI Overviews study and Profound's platform citation patterns are not measuring the same thing despite sounding like it.
Google's own documentation on AI features and its AI optimization guide describe both surfaces in one breath and distinguish neither. Publisher control is equally coarse: Google-Extended does not split by surface, so nobody can opt into one and out of the other.
The interface difference also changes what a citation is worth. An AI Overview sits above ordinary results, so an unsatisfying summary has a fallback one scroll away. A conversational surface has none. Referral volumes from it should not be expected to resemble classic organic, which is why conversion benchmarks matter more than visit counts. Open question Nobody has measured how citation patterns change across the turns of a conversation. We suspect the first answer is unrepresentative. We cannot show it.
Method: sample, conditions and schedule
Sample. The same 1,000-query set as the Citation Index, run through both surfaces inside a single collection window so that a product change cannot masquerade as a difference between surfaces.
Conditions held fixed. Logged out, one region, one device class, clean sessions, both surfaces queried the same day. Region, logged-in state and device all plausibly change the answer, so each is recorded on every observation rather than assumed away.
Unit of comparison. Cited domains per query, compared per query rather than pooled, with URL-level overlap recorded alongside because it is always the lower and more honest of the two.
Schedule. A first window establishes the baseline; the panel then repeats on the Citation Index cadence, because a single reading of two products that change without announcement has a short shelf life and an undated figure circulates forever. Conventions follow the measurement standard. It is the same method the cross-platform concordance work applies across companies, turned inward on one, and it is registered on the studies index.
Which queries can actually be compared?
Underneath the comparison sits a sampling problem: the two surfaces do not appear for the same queries. An AI Overview appears on some searches and not others — how often is itself contested, and it interacts with the click-through and zero-click behaviour Ahrefs has measured directly. AI Mode is something a user chooses to enter.
So the comparable set is smaller than the query set, and it is not randomly selected: queries where both surfaces respond are systematically different from queries where only one does. Handling that honestly means reporting three groups — both responded, one responded, neither responded — applying the overlap statistic only to the first, and treating the size of the other two as a finding in its own right. Most informal comparisons quietly drop the non-matching queries, which inflates apparent similarity because the queries both surfaces handle are the easy ones.
Control: how much does a surface agree with itself?
Any comparison of two systems needs to know how much each system disagrees with itself, or randomness gets reported as difference. The procedure is simple: run the same query through the same surface more than once, under the same conditions, and measure how much the source list changes. Do it for both surfaces, before any cross-surface number is calculated.
If a surface agrees with itself sixty percent of the time, a cross-surface agreement of fifty percent is barely a gap. If it agrees with itself ninety-five percent of the time, the same fifty percent is a large one. The self-agreement figures would be published alongside every cross-surface number. Any comparison that omits them is reporting a difference it cannot distinguish from noise — which includes several of the informal comparisons currently circulating.
Analysis: the four numbers we would report
One statistic cannot carry this comparison. Four can, and each answers a different practical question.
Ranking correspondence has the most external context to check against. Ahrefs found wide variation across engines, while Seer found an unusually tight match on one of them. Domain diversity carries the most practical weight, and most-cited domains gives the cross-platform baseline.
Every figure is broken out by query type. A narrow factual question has few authoritative sources and little room to diverge; a broad comparative one gives deeper fan-out room to reach further. Reported together these four say more than any composite score could. Blended into one they would say almost nothing, which is why we will not build one.
Pre-registered hypotheses
| # | Hypothesis | Predicted outcome |
|---|---|---|
| MD1 | AI Mode's cited domain set overlaps with AI Overview's at less than 50% for the same query | Expect to hold |
| MD2 | AI Mode cites a broader range of distinct domains per query than AI Overviews, consistent with deeper fan-out retrieval | Expect to hold |
| MD3 | The delta is larger for broad comparative queries than for narrow factual ones | Expect to hold |
AI Mode and AI Overviews are both Google products using fan-out, and coverage often treats them as interchangeable. Whether they actually cite the same sources for the same query has never been directly measured.
Share on XA worked hypothetical comparison
Here's what the comparison will look like once real data exists, using invented numbers only. Say the query is "how does a heat pump compare to a gas furnace for cold climates."
AI Overviews cites three sources for it: a manufacturer page, a government energy-efficiency site, and a review publication. AI Mode, for the identical query, cites six sources. One is the same government site. The other five are a forum thread, two comparison articles, and a regional utility company's guidance page.
One source is shared out of eight distinct ones combined. That works out to a Jaccard index of roughly 0.125 for this single query. That's well below the 50% threshold hypothesis MD1 predicts holds on average across the full query set. It also fits MD2's prediction: that AI Mode's retrieval surfaces a meaningfully broader domain set than AI Overviews' more compact snippet.
What would each possible outcome mean?
| Outcome | What it would mean | Who it changes things for |
|---|---|---|
| Near-total overlap | One measurement covers both surfaces. The industry's implicit assumption holds. | Everyone — the cost of tracking Google halves. |
| Partial overlap, wider AI Mode | AI Mode reaches further down the web. Our stated expectation. | Specialist and mid-size sites, most of all. |
| Low overlap both ways | Google's own surfaces disagree about as much as separate companies do. | Anyone reporting a single "Google AI visibility" figure. |
| Overlap varying by query type | The right question becomes "do they agree about my kind of question". | Every vertical differently. |
| Difference below the noise floor | No delta can be reported at all; the self-agreement figures are the finding. | Anyone quoting informal comparisons. |
The practical stake is the same in every row: tracking "Google AI visibility" as one blended metric, as many tools do, would hide meaningfully different performance on two products. That is the measurement failure behind incomparable headline numbers in this field — a problem even careful roundups inherit from their sources.
Open question Until the study runs, the defensible position is to split every Google AI figure by surface and date it. Then answer both narrowly and in depth on the same page. A clear top-line answer serves a fast retrieval pass — close to what Onely calls LLM-friendly content and what the citation playbook recommends. Explicit sub-answers underneath serve a fan-out pass. That hedge holds under every outcome above, which is the only reason to recommend it before the data exists. Freshness is another uncontrolled variable, common enough as a claim that we registered a test for it.
Who would a delta matter most to?
Large established publishers. If the surfaces converge on well-known sources, these sites win on both and the delta barely affects them. Their risk is not visibility, it is the click that no longer arrives.
Specialist and mid-size sites. Where the delta would matter most. A surface that reaches further down the web is the one where a specialist can appear at all, so if AI Mode is genuinely broader it is the more winnable surface and worth measuring separately.
New sites with no ranking history. If citations track classic rankings closely, a new site is invisible on both surfaces until it ranks — the exact question the zero-to-cited log study followed from a standing start. If they do not, there is a path in that does not require years of accumulated authority, which is a large difference in what a launch plan should look like.
B2B sellers. Their questions are narrow and often have no canonical source — the same gap that makes Perplexity's Reddit skew and ChatGPT's Wikipedia skew less binding for them than the headline numbers suggest. Fan-out into sub-questions plausibly helps them more than anyone. Plausibly, again, because nobody has measured it. Local businesses are a separate problem: both surfaces lean on different machinery for local intent, and this study says nothing about it.
The null results we would publish
Written before the data exists, so they cannot be quietly dropped later, and destined for the null results registry.
- Indistinguishable surfaces. We publish that, and retire the hypothesis that motivated the study.
- A real difference smaller than the noise floor. We publish the noise floor and decline to report a delta.
- Too few queries returning both surfaces. We publish the coverage numbers and say the study could not be run as designed.
- Fan-out prediction inverted. If AI Mode cites fewer sources rather than more, that goes on the page next to the prediction it contradicts. Being wrong in public is the entire value of pre-registering.
Reproduction: checking this yourself, cheaply
You do not need a thousand queries to learn something useful about your own topic. Twenty will do — and the tools here plus the Claude Code for SEO guide cover automating the tedious parts.
- 01 Fix twenty questions Questions your buyers actually ask. Write the list down and do not edit it.
- 02 Run both surfaces Same day, same place, logged out. Record which surfaces responded at all.
- 03 Record every source Domains are enough to start with.
- 04 Repeat five queries This gives you your own noise floor — the step that separates observation from coincidence.
- 05 Count three things Shared sources, sources unique to each surface, and sources per answer.
You will not get a publishable result. You will get an answer to the only version of the question that affects your budget, which is more than any industry average can offer you. Label the surface on every observation as you go: it costs nothing today and makes a year of records usable later.
Raw data and what gets published
When the first window closes, the intended release is the raw material rather than a summary. That means the query set, the per-query citation lists from each surface with timestamps and conditions, the three coverage groups, the repeat-run pairs behind the noise floor, and the four metrics per query type. Queries dropped for any reason are listed with the reason, so the denominator is visible and anyone can recompute the headline figures. None of it exists yet — there is no dataset to download at this stage, and this section describes what will be published, not what has been.
Limitations
- Sample and query selection. One shared 1,000-query set, and only the subset where both surfaces respond is comparable. That subset is not random — it skews toward the queries both products find easy — and a differently weighted set could move the headline overlap figure substantially.
- Geography and session state. Logged out, one region, one device class. Personalisation, locale routing and account history are uncontrolled, and results describe a fresh anonymous user rather than a typical one.
- Engine coverage. Google only. ChatGPT, Perplexity, Claude and Gemini behave differently again, as Leapd and Discovered Labs both report. Agent-facing surfaces such as agentic browsers and MCP endpoints are out of scope too.
- Measurement. Only the first answer in a conversational surface is captured; each follow-up turn is another retrieval and is not measured. Overlap also measures agreement, not which surface's citations are better — that needs a separate quality-evaluation method this study does not attempt.
- Confounders. Publisher controls are coarse and not split by surface, so opt-out state cannot be isolated. Page freshness, indexing recency and triggering changes can all shift citations independently of anything a site did.
- Reproducibility and generalisation. Both surfaces are actively developed and may converge or diverge after the window closes, so every figure is a dated snapshot rather than an architectural fact. Terminology and grading conventions are set out in how this site works.
In the meantime, the AI Overview exposure checker will tell you which of your own queries are likely to return an AI Overview at all — the surface this comparison treats as the baseline. The paired citation files and the per-surface breakdown go out to the newsletter when the first collection window closes — that is how to get notified when this reports. If you want to propose queries for the comparison set or point out a flaw in the design before it runs, the about page explains how to get in touch. There is no sign-up form; a message is enough.
Namdev, R. (2026). AI Mode vs. AI Overviews: source delta (v1). Retrieved from https://ritiknamdev.com/blog/ai-mode-vs-ai-overviews-source-delta Published under CC BY 4.0 — reuse freely with attribution.
Builds on AI Overview statistics and AI Mode statistics — read both for the mechanism behind each surface before this comparison.