Original research · Pre-registered

AI Mode vs. AI Overviews: source delta

Two distinct Google AI surfaces, routinely treated as interchangeable in coverage. Do they actually cite the same sources for the same queries? Nobody has published a direct comparison — this is the pre-registered protocol for making one.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — design stage ·13 min read ·Last verified September 2026
The short version

The question: do Google's two AI surfaces cite the same sources for the same query? Both are documented to use "query fan-out," but they have different interfaces and likely different retrieval depth, and no direct same-query comparison has been published by anyone. This page registers the query set, the four metrics and the predictions before a single paired query has been run.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. This is the within-Google concordance test. No paired queries have been run.

What is known
  • AI Mode and AI Overviews are separate products, and coverage routinely conflates them.
  • Both surfaces reportedly use query fan-out, so some shared retrieval behaviour is plausible.
What is not yet known
  • Whether the two surfaces cite the same sources for the same query, and how far they diverge.
  • Whether optimising for one surface transfers to the other.
  • Whether the delta is stable or moves as both products change.

What question is this study trying to answer?

Take an identical query set, run it through both surfaces under comparable conditions: what share of cited domains overlap, and where they diverge, what characterises the difference?

It is unanswered because nobody has published a direct, same-query comparison. Both surfaces belong to Google and both reportedly use fan-out — a mechanism Google described in some detail for AI Mode specifically. So most coverage treats them as interchangeable, differing only in placement: embedded snippet versus dedicated mode. That conflation is one of the recurring errors documented in the statistics-provenance audit. It has to be got right before any Google figure means anything.

The comparison also only makes sense if both surfaces can reach the page at all, which puts crawler access and robots.txt policy upstream of everything here.

What is known about each surface?

Evidence

AI Overview citation-to-ranking overlap fell from ~76% to ~38% year over year, per Ahrefs' top-10 citation analysis, reported by SEJ as a sharp drop and documented in the dedicated statistics page.

Open question

AI Mode's comparable overlap figure, and how it compares directly to AI Overviews' on the same query set, hasn't been published.

~76% → ~38%

AI Overview citations that also ranked in Google's top 10, year over year. The surface moved away from conventional rankings fast.

Ahrefs
Unmeasured

AI Mode's comparable overlap figure, on the same query set, in the same window. No public comparison exists.

Open question

Why would two Google surfaces cite differently?

AI OverviewsAI Mode
Where it livesA box above conventional resultsA dedicated conversational surface
How a user reaches itAppears, unprompted, on some queriesEntered deliberately
Fallback in viewThe classic listings, right belowNone — the answer is the page
Retrieval depth (reported)Narrow, fast passExtensive fan-out across sub-queries
Follow-up turnsNot part of the surfaceEach turn is another retrieval
Publisher controlOne coarse AI opt-out, not split by surface

Fan-out means a single question becomes several searches. The system breaks the question into narrower parts, searches for each, and assembles one answer from everything it gathered. "Is a standing desk worth it for lower back pain?" might become separate searches for health effects, back pain causes, sit-stand alternation evidence, and typical costs — four searches, four source pools.

A surface that runs closer to one search and summarises the top results draws from a much narrower pool, one that overlaps heavily with what already ranks. That is the crux. If one surface fans out and the other does not, they are not two skins on one system but two retrieval processes that share a company. It is also why Semrush's AI Overviews study and Profound's platform citation patterns are not measuring the same thing despite sounding like it.

Google's own documentation on AI features and its AI optimization guide describe both surfaces in one breath and distinguish neither. Publisher control is equally coarse: Google-Extended does not split by surface, so nobody can opt into one and out of the other.

The interface difference also changes what a citation is worth. An AI Overview sits above ordinary results, so an unsatisfying summary has a fallback one scroll away. A conversational surface has none. Referral volumes from it should not be expected to resemble classic organic, which is why conversion benchmarks matter more than visit counts. Open question Nobody has measured how citation patterns change across the turns of a conversation. We suspect the first answer is unrepresentative. We cannot show it.

Method: sample, conditions and schedule

Sample. The same 1,000-query set as the Citation Index, run through both surfaces inside a single collection window so that a product change cannot masquerade as a difference between surfaces.

Conditions held fixed. Logged out, one region, one device class, clean sessions, both surfaces queried the same day. Region, logged-in state and device all plausibly change the answer, so each is recorded on every observation rather than assumed away.

Unit of comparison. Cited domains per query, compared per query rather than pooled, with URL-level overlap recorded alongside because it is always the lower and more honest of the two.

Schedule. A first window establishes the baseline; the panel then repeats on the Citation Index cadence, because a single reading of two products that change without announcement has a short shelf life and an undated figure circulates forever. Conventions follow the measurement standard. It is the same method the cross-platform concordance work applies across companies, turned inward on one, and it is registered on the studies index.

Which queries can actually be compared?

Underneath the comparison sits a sampling problem: the two surfaces do not appear for the same queries. An AI Overview appears on some searches and not others — how often is itself contested, and it interacts with the click-through and zero-click behaviour Ahrefs has measured directly. AI Mode is something a user chooses to enter.

So the comparable set is smaller than the query set, and it is not randomly selected: queries where both surfaces respond are systematically different from queries where only one does. Handling that honestly means reporting three groups — both responded, one responded, neither responded — applying the overlap statistic only to the first, and treating the size of the other two as a finding in its own right. Most informal comparisons quietly drop the non-matching queries, which inflates apparent similarity because the queries both surfaces handle are the easy ones.

Control: how much does a surface agree with itself?

Any comparison of two systems needs to know how much each system disagrees with itself, or randomness gets reported as difference. The procedure is simple: run the same query through the same surface more than once, under the same conditions, and measure how much the source list changes. Do it for both surfaces, before any cross-surface number is calculated.

If a surface agrees with itself sixty percent of the time, a cross-surface agreement of fifty percent is barely a gap. If it agrees with itself ninety-five percent of the time, the same fifty percent is a large one. The self-agreement figures would be published alongside every cross-surface number. Any comparison that omits them is reporting a difference it cannot distinguish from noise — which includes several of the informal comparisons currently circulating.

Analysis: the four numbers we would report

One statistic cannot carry this comparison. Four can, and each answers a different practical question.

Source overlapOf the sources cited by both surfaces, what fraction is shared? Answers whether visibility transfers.
Source countDistinct sources per answer. Tests the fan-out prediction and corrects overlap for verbosity.
Ranking correspondenceShare of cited pages also in the classic top ten. Says how much existing SEO position carries over.
Domain diversityHow concentrated each surface is on large publishers. The number that decides whether a small site can win.

Ranking correspondence has the most external context to check against. Ahrefs found wide variation across engines, while Seer found an unusually tight match on one of them. Domain diversity carries the most practical weight, and most-cited domains gives the cross-platform baseline.

Every figure is broken out by query type. A narrow factual question has few authoritative sources and little room to diverge; a broad comparative one gives deeper fan-out room to reach further. Reported together these four say more than any composite score could. Blended into one they would say almost nothing, which is why we will not build one.

Pre-registered hypotheses

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
MD1AI Mode's cited domain set overlaps with AI Overview's at less than 50% for the same queryExpect to hold
MD2AI Mode cites a broader range of distinct domains per query than AI Overviews, consistent with deeper fan-out retrievalExpect to hold
MD3The delta is larger for broad comparative queries than for narrow factual onesExpect to hold

AI Mode and AI Overviews are both Google products using fan-out, and coverage often treats them as interchangeable. Whether they actually cite the same sources for the same query has never been directly measured.

Share on X

A worked hypothetical comparison

Here's what the comparison will look like once real data exists, using invented numbers only. Say the query is "how does a heat pump compare to a gas furnace for cold climates."

AI Overviews cites three sources for it: a manufacturer page, a government energy-efficiency site, and a review publication. AI Mode, for the identical query, cites six sources. One is the same government site. The other five are a forum thread, two comparison articles, and a regional utility company's guidance page.

One source is shared out of eight distinct ones combined. That works out to a Jaccard index of roughly 0.125 for this single query. That's well below the 50% threshold hypothesis MD1 predicts holds on average across the full query set. It also fits MD2's prediction: that AI Mode's retrieval surfaces a meaningfully broader domain set than AI Overviews' more compact snippet.

What would each possible outcome mean?

OutcomeWhat it would meanWho it changes things for
Near-total overlapOne measurement covers both surfaces. The industry's implicit assumption holds.Everyone — the cost of tracking Google halves.
Partial overlap, wider AI ModeAI Mode reaches further down the web. Our stated expectation.Specialist and mid-size sites, most of all.
Low overlap both waysGoogle's own surfaces disagree about as much as separate companies do.Anyone reporting a single "Google AI visibility" figure.
Overlap varying by query typeThe right question becomes "do they agree about my kind of question".Every vertical differently.
Difference below the noise floorNo delta can be reported at all; the self-agreement figures are the finding.Anyone quoting informal comparisons.

The practical stake is the same in every row: tracking "Google AI visibility" as one blended metric, as many tools do, would hide meaningfully different performance on two products. That is the measurement failure behind incomparable headline numbers in this field — a problem even careful roundups inherit from their sources.

Open question Until the study runs, the defensible position is to split every Google AI figure by surface and date it. Then answer both narrowly and in depth on the same page. A clear top-line answer serves a fast retrieval pass — close to what Onely calls LLM-friendly content and what the citation playbook recommends. Explicit sub-answers underneath serve a fan-out pass. That hedge holds under every outcome above, which is the only reason to recommend it before the data exists. Freshness is another uncontrolled variable, common enough as a claim that we registered a test for it.

Who would a delta matter most to?

Large established publishers. If the surfaces converge on well-known sources, these sites win on both and the delta barely affects them. Their risk is not visibility, it is the click that no longer arrives.

Specialist and mid-size sites. Where the delta would matter most. A surface that reaches further down the web is the one where a specialist can appear at all, so if AI Mode is genuinely broader it is the more winnable surface and worth measuring separately.

New sites with no ranking history. If citations track classic rankings closely, a new site is invisible on both surfaces until it ranks — the exact question the zero-to-cited log study followed from a standing start. If they do not, there is a path in that does not require years of accumulated authority, which is a large difference in what a launch plan should look like.

B2B sellers. Their questions are narrow and often have no canonical source — the same gap that makes Perplexity's Reddit skew and ChatGPT's Wikipedia skew less binding for them than the headline numbers suggest. Fan-out into sub-questions plausibly helps them more than anyone. Plausibly, again, because nobody has measured it. Local businesses are a separate problem: both surfaces lean on different machinery for local intent, and this study says nothing about it.

The null results we would publish

Written before the data exists, so they cannot be quietly dropped later, and destined for the null results registry.

  • Indistinguishable surfaces. We publish that, and retire the hypothesis that motivated the study.
  • A real difference smaller than the noise floor. We publish the noise floor and decline to report a delta.
  • Too few queries returning both surfaces. We publish the coverage numbers and say the study could not be run as designed.
  • Fan-out prediction inverted. If AI Mode cites fewer sources rather than more, that goes on the page next to the prediction it contradicts. Being wrong in public is the entire value of pre-registering.

Reproduction: checking this yourself, cheaply

You do not need a thousand queries to learn something useful about your own topic. Twenty will do — and the tools here plus the Claude Code for SEO guide cover automating the tedious parts.

A twenty-query surface comparison you can run yourself
  1. 01 Fix twenty questions Questions your buyers actually ask. Write the list down and do not edit it.
  2. 02 Run both surfaces Same day, same place, logged out. Record which surfaces responded at all.
  3. 03 Record every source Domains are enough to start with.
  4. 04 Repeat five queries This gives you your own noise floor — the step that separates observation from coincidence.
  5. 05 Count three things Shared sources, sources unique to each surface, and sources per answer.

You will not get a publishable result. You will get an answer to the only version of the question that affects your budget, which is more than any industry average can offer you. Label the surface on every observation as you go: it costs nothing today and makes a year of records usable later.

Raw data and what gets published

When the first window closes, the intended release is the raw material rather than a summary. That means the query set, the per-query citation lists from each surface with timestamps and conditions, the three coverage groups, the repeat-run pairs behind the noise floor, and the four metrics per query type. Queries dropped for any reason are listed with the reason, so the denominator is visible and anyone can recompute the headline figures. None of it exists yet — there is no dataset to download at this stage, and this section describes what will be published, not what has been.

Limitations

  • Sample and query selection. One shared 1,000-query set, and only the subset where both surfaces respond is comparable. That subset is not random — it skews toward the queries both products find easy — and a differently weighted set could move the headline overlap figure substantially.
  • Geography and session state. Logged out, one region, one device class. Personalisation, locale routing and account history are uncontrolled, and results describe a fresh anonymous user rather than a typical one.
  • Engine coverage. Google only. ChatGPT, Perplexity, Claude and Gemini behave differently again, as Leapd and Discovered Labs both report. Agent-facing surfaces such as agentic browsers and MCP endpoints are out of scope too.
  • Measurement. Only the first answer in a conversational surface is captured; each follow-up turn is another retrieval and is not measured. Overlap also measures agreement, not which surface's citations are better — that needs a separate quality-evaluation method this study does not attempt.
  • Confounders. Publisher controls are coarse and not split by surface, so opt-out state cannot be isolated. Page freshness, indexing recency and triggering changes can all shift citations independently of anything a site did.
  • Reproducibility and generalisation. Both surfaces are actively developed and may converge or diverge after the window closes, so every figure is a dated snapshot rather than an architectural fact. Terminology and grading conventions are set out in how this site works.
When this runs

In the meantime, the AI Overview exposure checker will tell you which of your own queries are likely to return an AI Overview at all — the surface this comparison treats as the baseline. The paired citation files and the per-surface breakdown go out to the newsletter when the first collection window closes — that is how to get notified when this reports. If you want to propose queries for the comparison set or point out a flaw in the design before it runs, the about page explains how to get in touch. There is no sign-up form; a message is enough.

How to cite this
Namdev, R. (2026). AI Mode vs. AI Overviews: source delta (v1). Retrieved from https://ritiknamdev.com/blog/ai-mode-vs-ai-overviews-source-delta

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Builds on AI Overview statistics and AI Mode statistics — read both for the mechanism behind each surface before this comparison.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Search Engine Journal — Query fan-out in AI Mode: new details from Googlewww.searchenginejournal.com/query-fan-out-technique-in-ai-mode-new-details-from-google/552532 Google Search Central — AI features and your websitedevelopers.google.com/search/docs/appearance/ai-features Google Search Central — AI optimization guidedevelopers.google.com/search/docs/fundamentals/ai-optimization-guide Google Search Central — Overview of Google crawlers and fetchersdevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers AmICited — Google-Extended: what it does and whether to block itwww.amicited.com/blog/google-extended-what-it-does-should-you-block-it Ahrefs — AI Overview citations and top-10 rankingsahrefs.com/blog/ai-overview-citations-top-10 Ahrefs — AI Overviews reduce clicks (update)ahrefs.com/blog/ai-overviews-reduce-clicks-update Ahrefs — AI search overlap between platformsahrefs.com/blog/ai-search-overlap Ahrefs — AI SEO statisticsahrefs.com/blog/ai-seo-statistics Search Engine Journal — AI Overview citations from top-ranking pages drop sharplywww.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637 Semrush — AI Overviews studywww.semrush.com/blog/semrush-ai-overviews-study Similarweb — generative AI usage statisticsaisearch.similarweb.com/blog/gen-ai-stats StatCounter — Search engine market share worldwidegs.statcounter.com/search-engine-market-share Discovered Labs — How each platform cites sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Discovered Labs — AI citation patterns across platformsdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Leapd — How ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors Onely — What makes content LLM-friendlywww.onely.com/blog/llm-friendly-content Salespeak — Content freshness in AI searchsalespeak.ai/aeo-news/content-freshness-ai-search Wikipedia — Generative engine optimizationen.wikipedia.org/wiki/Generative_engine_optimization
FAQ

Frequently asked questions

Aren't AI Overviews and AI Mode basically the same thing with different UI?
That's the working assumption in a lot of coverage. It's exactly the assumption this study tests, rather than accepts. AI Mode's documented query fan-out mechanism suggests a meaningfully different retrieval process than AI Overviews' embedded-snippet approach. That's true even though both surfaces sit on shared Google infrastructure.
How would you even run identical queries through both surfaces fairly?
By using the same query set, from a comparable session state, logged out, consistent geography, for both surfaces. Citations from each get logged independently. It's the same core method as the rest of the Citation Index, applied specifically to tell these two Google products apart.
If the two surfaces do diverge, does that mean one is "better" than the other?
Not necessarily. Divergence would say the two surfaces retrieve and cite differently. It wouldn't say one set of citations is higher quality than the other. This study measures overlap and breadth, not citation quality. That would need a separate evaluation method entirely.
Why use the Citation Index's existing 1,000-query set rather than a purpose-built one?
Reusing the shared query set means this comparison needs no separate query-design work. It also stays directly comparable to every other citation measurement built on the same set. A query set built just for this comparison would add design overhead with no clear analytical benefit.
Can I control whether I appear in one surface but not the other?
No. Google offers a control over AI use of your content that is separate from Search indexing, but it does not split by surface. There is no way to allow AI Overviews and refuse AI Mode. That coarseness is one of the reasons these two products are hard to study independently.
Why not just compare screenshots people have posted?
Because they were captured at different times, in different regions, from different account states, on different queries. Every one of those varies the result. A comparison is only meaningful when the conditions are held fixed on purpose.
If the two surfaces cite nearly the same sources, is this study wasted?
No, that is a genuinely useful result. It would mean one measurement covers both surfaces, which halves the ongoing cost of tracking Google. We would publish it as clearly as we would publish a divergence.
Does appearing in AI Overviews send traffic?
Sometimes, and less than the equivalent classic result usually did. The surface is designed to answer without a click. Treat any appearance as brand exposure first and traffic second, and measure both separately.
What should I report to a client before this study runs?
Split every Google AI figure by surface, date every observation, record region and logged-in state, and report an absence as unknown rather than zero. A combined "Google AI visibility" number is indefensible today, and splitting it costs nothing.
Why pre-register the design instead of publishing once the data is in?
A design written after the fact can be reshaped to fit whatever the numbers said. Registering the query set, the metrics and the predictions first means anyone can check the goalposts did not move — including on the predictions we get wrong.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.