Original research · Pre-registered · Emerging surface

What Agentic Browsers Actually Fetch

ChatGPT Atlas, Perplexity Comet, and Gemini in Chrome are shipping now. What they actually request from a page — and whether it differs from a standard bot fetch — has barely been studied. This is the pre-registered protocol for measuring it.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — instrumentation stage ·13 min read ·Last verified September 2026
The short version

The question: when an agentic browser opens your page, does it fetch like a crawler, render like a browser, or click like a human? Nobody has published a controlled answer. This page registers the instrumentation and the predictions before any session is recorded. Practitioner write-ups already treat agentic traffic as a category to prepare for, without a measurement behind the advice.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. Instrumentation is being built. No agentic-browser session has been recorded under this protocol.

What is known
  • Agentic browsers identify themselves with documented user-agent strings, catalogued in the bot registry.
  • A 30-day site-log study of agentic crawler behaviour exists as a methodological precedent.
What is not yet known
  • What an agentic browser actually fetches versus a classic crawler.
  • Whether it executes JavaScript, and whether it respects robots.txt in practice.
  • How much of a page it retrieves before acting.

What question is this study trying to answer?

Does an agentic browser's fetch behave like a standard retrieval bot, like a full human-driven browser session, or like something in between? Does it respect robots.txt for the specific fetch it performs? Does it execute JavaScript — and beyond that, does it click, expand and scroll to reach content a human would have to click for?

Those questions are unanswered because the surface is new and the operators have not documented it. ChatGPT Atlas, Perplexity Comet and Gemini's in-Chrome assistant are all live products, while every published bot document from those same operators describes the autonomous crawler instead. That gap is the whole reason for this protocol: measure the behaviour before it gets assumed and repeated as folklore, the failure mode catalogued in where AI SEO statistics come from. If the vocabulary is unfamiliar, start with the AI search glossary and what GEO is.

What is already known, and what is not

Fact

Each of these products can navigate, summarize, and take actions on a live web page at a user's direction, distinct from a chatbot's independent web search tool.

Open question

Whether they render JavaScript, respect robots.txt for the specific fetch they perform, or leave a distinguishable trace in server logs, hasn't been rigorously documented in public. The nearest thing is a 30-day single-site log study, and it does not separate the four steps below.

ProductOperatorInteraction model
ChatGPT AtlasOpenAIDedicated agentic browser
CometPerplexityDedicated agentic browser
Gemini in ChromeGoogleIn-browser assistant mode

Each operator publishes bot documentation for its autonomous crawlers — OpenAI, Perplexity, Anthropic and Google — and none of it describes the agentic-browser fetch path. The AI Bot Registry collects what is documented; AI crawler statistics covers the volumes.

Why is this not the same question as AI crawling?

A retrieval or training bot's incentives favour speed and scale: fetch as many pages as cheaply as possible, generally without executing JavaScript, because running a browser engine across billions of pages does not pencil out. That cost is measured — Google needs roughly nine times longer to crawl JavaScript than HTML, with a median rendering delay of several seconds.

An agentic browser has the opposite incentive structure. It renders one page, at one moment, for one user actively waiting. That resource profile is closer to a real browser tab — the page-weight and interaction budget the Web Almanac performance chapter and the Core Web Vitals thresholds describe — than to a batch crawler. So the protocol does not assume agentic browsers behave like crawlers, even when one company runs both. Anthropic, for instance, documents an answer-time web search tool entirely separately from its training crawler.

DimensionAutonomous crawlerAgentic browser
TriggerA schedule or a queueOne user, one request, right now
ScaleBillions of pagesOne page at a time
JavaScriptUsually skipped for cost reasonsPlausibly executed — untested
InteractionNonePossible: clicks, tabs, scrolling
robots.txtDocumented per botUndocumented for a user-directed fetch
Log signatureDeclared user agentUnknown, possibly indistinguishable from a human

What happens when an agent opens a page?

The dependent variables come from separating four steps that public commentary collapses into one question: "does it read my page?" That question cannot be answered usefully. The four steps fail in different ways, and each has a different fix.

The four steps an agentic browser takes, and where each one fails
  1. 1 Fetch A request with a user agent and headers. Failure mode: the site cannot tell this is not a person.
  2. 2 Render Raw HTML parsed as text, or a full browser engine running the scripts. Failure mode: client-side content is invisible.
  3. 3 Interact Clicking a tab, expanding an accordion, scrolling for lazy content. Failure mode: content behind a click is never reached.
  4. 4 Extract Deciding what answers the user. Failure mode: the answer sits below three screens of preamble.

Open question Whether these products identify themselves consistently at step one is not documented publicly. Step three — interaction — is the step that separates an agent from a crawler, and the least studied of the four.

Method: sample, control, variables and schedule

Sample. One instrumented test site under our control, one page per test condition, three products: Atlas, Comet and Gemini in Chrome. Each product is run repeatedly across several days rather than once, because a single run is a snapshot rather than a behaviour.

Control. The same page visited by an ordinary human-driven browser session, logged identically, establishes what a normal visit looks like in our own logs. Without that baseline there is nothing to call an agentic fetch unusual against. A second reference point is the classic-crawler baseline from the JavaScript rendering experiment on the same markup.

Independent variable. The client: human browser, classic retrieval bot, or each agentic browser in turn, with the prompt held fixed and written down.

Dependent variables. Which of three planted strings the client can report back — one in the static HTML, one revealed only after JavaScript runs, one revealed only after a button is clicked — plus the server-side record: full user agent, headers, IP, timestamp, path, referer, and whether robots.txt was requested at all.

Analysis. Per product, never averaged across products. Each run is scored on the highest of the four steps it demonstrably reached, with counts reported as counts, not rates, until the sample is large enough to support a rate. Where server evidence and the agent's own narration disagree, the log wins and the disagreement is recorded as its own observation.

Schedule. Instrumentation first: the test page, the server-side logging and the human control session. Then a first collection window in which every product is run against every condition inside a short span of days, so a mid-window product release cannot be mistaken for a difference between products. The window is deliberately narrow for that reason, which is also why the sample stays small.

The panel then repeats on a fixed cadence rather than running once. A single reading of a product that ships weekly has a short shelf life, and a number with no expiry date circulates forever. Each repeat re-runs the identical page and prompts, so the comparison between windows — not any single window — is where the information lives.

The procedure follows the AI visibility measurement standard, is recorded under the dataset strategy, and is registered on the studies index alongside the rest of the AI Citation Index roadmap. The three-string design is the one used in zero to cited, extended by one level to cover interaction.

Pre-registered hypotheses

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
AB1Agentic browsers render JavaScript on the pages they visit, unlike autonomous retrieval botsExpect to hold
AB2Agentic browsers respect robots.txt directives less consistently than autonomous crawlers do, given their user-directed rather than bulk-fetch natureExpect to hold
AB3Agent self-reports of what they did on a page diverge from the server-side evidence in at least some runsExpect to hold

ChatGPT Atlas, Perplexity Comet, and Gemini in Chrome are already shipping. What they actually fetch from a page — and whether it behaves like a bot or like a human browser — has barely been studied.

Share on X

What could make a result look real when it is not?

Several things can produce a pattern that looks like agent behaviour and is not. Naming them in advance is part of the method.

Caching and prefetch. A page the agent already holds may never be re-fetched, so an absent log line is not proof it ignored you. Conversely, some fetches happen before a user commits to a page; those look identical to a real read, but nobody saw the result.

Shared infrastructure. A request may come from a general fetching service used by several features, so attributing it to the agentic browser specifically may simply be wrong.

CDN and edge behaviour. Your edge may serve a cached variant, strip a header, or challenge the request — in which case origin logs tell a story about your CDN, not the agent. Blanket AI-blocking rules are easy to apply by accident, as large-scale robots.txt analysis and a dedicated blocking report both show. Our own count is the robots.txt AI-blocking census, and GPTBot vs OAI-SearchBot covers the distinction people most often get wrong when writing one.

Version drift and prompt sensitivity. A product may update mid-window, so a change in the data may be a release rather than a finding. And how the agent is asked to visit changes what it does: "summarize this page" and "find the hidden detail" are different experiments, both valid, never to be mixed.

The test page itself. A page built purely for testing has no inbound links, no history and unusual markup, and an agent may treat it differently from a real page. Every observation therefore records the date, the visible product version, and the serving path.

Null results we would publish

  • No distinguishable trace. If agentic browser visits are indistinguishable from ordinary human sessions in server logs, that is a significant, publishable result.
  • No interaction at all. If no product clicks the button, the interaction hypothesis fails and the "agents behave like humans" framing weakens.
  • No difference from the crawler baseline. If agentic fetches match ordinary retrieval bots, the case for treating this as a separate surface collapses.
  • Unstable results. If repeated runs disagree with each other, we report the instability rather than picking the run that reads best.

Refusals count too: if an agent declines to visit, or answers something unrelated, that is recorded and filed in the null results registry rather than dropped.

What would a result actually mean for a site?

The following example is invented to illustrate the mechanism. It is not a case study, and no client or measured result is being described. Take a hardware product page: the price sits in the initial HTML, the specifications sit inside a tab rendered only after a click, the warranty terms sit in an accordion. An agent asked to compare that product against two rivals can read the price and may never reach the specifications, so the answer the user receives is thin and possibly wrong by omission.

Move the specifications into the initial HTML as a plain table and nothing changes for the human; everything changes for a client that does not click. That interaction cost is not hypothetical for people either — it is why interaction latency became a Core Web Vital.

Hypothesis The prediction is that the second version is legible to more agents than the first. That prediction has not been tested here, and this page will not claim it has.

The implied fix is not detecting the agent, and not serving it a special page. It is making one page that works for a client with fewer capabilities — the same move that survived the AJAX era, and the principle underneath the technical GEO audit.

Nothing here is a ranking factor. This measures fetch and interaction behaviour. Whether an agent chooses to cite you is a separate question, covered by the cross-platform concordance protocol. Adjacent structured interfaces — MCP servers and WebMCP, placed in context by a survey of the 2026 agent-protocol stack — would change the question from "what can it render" to "what do you expose", which is why the baseline is worth taking now.

Reproduction: running this yourself

You do not need a lab. You need one page, a server log, and patience. The protocol is deliberately cheap enough to replicate.

String one — static HTMLPresent in the raw response. If the agent misses this, it never fetched the page you think it did.
String two — after JavaScriptAppears only once scripts run. Reaching it proves rendering, not interaction.
String three — after a clickThe only proof of genuine interaction, and the bar no crawler study tests.
The agent's own narrationA claim to be checked against the server log, never evidence on its own.

Build the page with those three strings. Log the full user agent, IP, timestamp, path and referer at the server rather than in client-side analytics. Visit once with a normal browser to establish your control.

Then direct each agent to the page one at a time, using the same written-down prompt, noting the date and any visible product version, and ask which strings it can see. Repeat across several days. Published user-agent inventories — this crawler reference, plus OpenAI's and Perplexity's own docs — are the right starting grep list, even though none of them names an agentic-browser agent explicitly. The free tools here cover the first checks, and Claude Code for SEO covers running the whole loop on a schedule.

Raw data and what gets published

When the first collection window closes, the release is raw material rather than a summary. That means the test page source, the exact prompts used, the per-run server log lines with timestamps and user agents, the agent transcripts, and a per-product table of which of the four steps each run reached. Runs that failed or were discarded are listed with the reason, so the denominator is visible. Nothing exists yet — there is no dataset to download at this stage, and this section describes what will be published, not what has been.

Open questions

  • Does a robots.txt disallow apply to a fetch a user explicitly requested? Different products may answer this differently, and none has published a clear rule.
  • Do agents follow internal links, or stop at the page they were given?
  • Does an agent see the same page a logged-out human sees, or does it inherit the user's session?
  • How does an agent handle a consent banner, a paywall, or an age gate?
  • Is there any relationship between what an agent fetches and what the same company's chatbot later cites? We have no evidence either way.

Limitations

  • Sample. One test site, three products, a handful of runs per product. Counts at that scale cannot support a percentage, and this study will not report one.
  • Engine coverage. Atlas, Comet and Gemini in Chrome only. Other assistants' browsing modes, enterprise agents and headless automation frameworks are outside the panel, and results will not be extrapolated to them — behaviour does not transfer between products, even between two products from the same operator.
  • Page selection. A purpose-built test page with no inbound links, no history and deliberately unusual markup is not a representative page. Behaviour observed on it may not match behaviour on a real, established URL.
  • Geography and session. All runs originate from a single region on a single network, from fresh accounts. Personalisation, locale routing and account history are uncontrolled and may change what an agent does.
  • Measurement. Server logs record a fetch, not an intention. Caching, prefetch, shared fetching infrastructure and CDN behaviour can each hide or fabricate a request, and an agent's own narration is never treated as evidence.
  • Reproducibility. These products ship changes weekly, so an identical re-run months later may legitimately produce different results. Every observation is dated and versioned for that reason; a mid-window product release is a confound we can note but not remove.
  • Generalisation. Volume is unmeasured — no operator breaks out agentic browser traffic, so nothing here can be weighted by how much of it exists. The figure sets in AI SEO statistics and the state of AI search have the same gap.
Where to go next

To find out what your own pages expose to a client that does not click: plant the three test strings described above and read your server logs, then work through the technical GEO audit for the rest of the surface. To see which bots you are already allowing or blocking: check the identifiers in the AI bot user-agent registry.

When this runs

Results, the test page source and the raw log files go out to the newsletter when the first collection window closes — that is the way to get notified when this reports. If you want to propose a test condition, offer a site to instrument, or point out a flaw in the design before it runs, the about page explains how to get in touch. There is no sign-up form for participation; a message is enough.

How to cite this
Namdev, R. (2026). What Agentic Browsers Actually Fetch (v1). Retrieved from https://ritiknamdev.com/blog/agentic-browsers-what-they-fetch

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Complements the JavaScript rendering experiment and the AI Bot Registry, which covers autonomous crawlers rather than user-driven agentic browsing.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Agentlux — Agentic traffic is here: how websites should prepare for AI browsers and shopping agentsagentlux.ai/blog/agentic-traffic-is-here-how-websites-should-prepare-for-ai-browsers-and-shopping-agents Previsible — Agentic shoppingprevisible.io/seo-ai-news/agentic-shopping Digital Applied — Agentic crawler behaviour: a 30-day site log studywww.digitalapplied.com/blog/agentic-crawler-behavior-30-day-site-log-study Dev.to — The state of agentic AI standards in 2026: MCP, A2A, WebMCP and the protocol stackdev.to/alexmercedcoder/the-state-of-agentic-ai-standards-in-2026-mcp-a2a-webmcp-osi-and-the-protocol-stack-taking-3o2l OpenAI — Bots and crawler documentationplatform.openai.com/docs/bots OpenAI — Bots reference (developer docs)developers.openai.com/api/docs/bots Perplexity — Crawler documentationdocs.perplexity.ai/docs/resources/perplexity-crawlers Perplexity — Developer documentationdocs.perplexity.ai Anthropic — Documentation homedocs.anthropic.com Anthropic — Web search tool for agentsplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Anthropic Support — Does Anthropic crawl the web, and how site owners can block the crawlersupport.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Cloudflare — Perplexity is using stealth, undeclared crawlers to evade no-crawl directivesblog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives Google Search Central — Overview of Google crawlers and fetchersdevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers Onely — Google rendering delay is about five secondswww.onely.com/blog/googles-rendering-delay-5-seconds Onely — Google needs 9x more time to crawl JavaScript than HTMLwww.onely.com/blog/google-needs-9x-more-time-to-crawl-js-than-html Web Almanac 2025 — Performance chapteralmanac.httparchive.org/en/2025/performance web.dev — Defining the Core Web Vitals thresholdsweb.dev/articles/defining-core-web-vitals-thresholds web.dev — INP becomes a Core Web Vitalweb.dev/blog/inp-cwv-march-12 Paul Calvano — AI bots and robots.txtpaulcalvano.com/2025-08-21-ai-bots-and-robots-txt Technology Checker — robots.txt AI crawler blocking reporttechnologychecker.io/blog/robots-txt-ai-crawlers-blocking-report Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots
FAQ

Frequently asked questions

What is an agentic browser, exactly?
A browser, or browser mode, with an AI agent embedded that can navigate, read, summarize, and take actions on web pages on a user's behalf. ChatGPT Atlas, Perplexity Comet, and Gemini in Chrome are current examples. They differ from a chatbot's web search tool because they interact with a live page, not a search index.
Isn't this the same as the existing AI crawler question?
Related, but distinct. Training and retrieval bots fetch pages autonomously at scale, to build an index or answer a query. Agentic browsers act on a page a specific user is currently pointed at, often executing JavaScript and interacting with page elements the way a human would. That's a genuinely different technical surface.
Would an agentic browser visit be logged differently from a normal human visit in analytics?
Potentially, if the product identifies itself distinctly in its request headers or user-agent string. But whether it does so consistently, or presents as indistinguishable from a normal browser session, is itself one of the open questions this study is designed to document, not assume.
Could a site deliberately serve different content to an agentic browser than to a human?
Technically possible, if the agent is reliably identifiable, similar in spirit to dynamic rendering for classic bots. But whether that's advisable, and whether these products' own policies permit it, is a separate and more ethically loaded question. This protocol doesn't take a position on it. It aims first to establish what actually happens under normal, unmodified serving.
Should I block agentic browsers in robots.txt?
There is no evidence base yet to answer that well. A block may remove you from an answer a real user asked for, because the fetch is user-directed rather than a bulk crawl. It may also protect a paywall. The honest position is that the trade-off has not been measured in public, so treat any confident recommendation you read with suspicion.
Do agentic browsers show up in Google Analytics?
Unknown, and it depends on whether the agent executes the analytics script at all. A fetch that skips JavaScript leaves no client-side analytics record but still appears in server logs. That gap is one reason this study prefers raw server logs over an analytics dashboard.
If an agent renders JavaScript, does that mean my client-side content is safe?
Not necessarily. Rendering on load and interacting with a page are different capabilities. Content behind a click, a tab, an accordion or an infinite scroll may still be invisible even to an agent that runs scripts. The protocol separates these two levels deliberately.
How many agentic browser visits does a normal site get?
No credible public figure exists for this, and we will not invent one. Traffic share for these products is not broken out by any operator in a form that a site owner can verify. The only number you can trust is the one in your own server logs.
Will the results of this study still be true in a year?
Probably not in detail. These products ship changes weekly. What should survive is the measurement method, and the categories of behavior worth watching. Treat any eventual finding as a dated snapshot rather than a standing rule.
Is there a quick check I can run in five minutes?
Yes. Open your most important page, disable JavaScript in the browser, and read what remains. Then re-enable it and note everything that only appears after you click something. Those two lists are a rough map of what a weaker agent will miss.
Why publish the protocol before running it?
Because a design written afterwards can be quietly reshaped to fit whatever the logs turned out to say. Registering the test page, the prompts and the predictions first is the cheapest protection against that, and it lets anyone check the goalposts did not move.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.