The question: when an agentic browser opens your page, does it fetch like a crawler, render like a browser, or click like a human? Nobody has published a controlled answer. This page registers the instrumentation and the predictions before any session is recorded. Practitioner write-ups already treat agentic traffic as a category to prepare for, without a measurement behind the advice.
No data has been collected yet. Nothing on this page is a result. Instrumentation is being built. No agentic-browser session has been recorded under this protocol.
- Agentic browsers identify themselves with documented user-agent strings, catalogued in the bot registry.
- A 30-day site-log study of agentic crawler behaviour exists as a methodological precedent.
- What an agentic browser actually fetches versus a classic crawler.
- Whether it executes JavaScript, and whether it respects robots.txt in practice.
- How much of a page it retrieves before acting.
What question is this study trying to answer?
Does an agentic browser's fetch behave like a standard retrieval bot, like a full human-driven browser session, or like something in between? Does it respect robots.txt for the specific fetch it performs? Does it execute JavaScript — and beyond that, does it click, expand and scroll to reach content a human would have to click for?
Those questions are unanswered because the surface is new and the operators have not documented it. ChatGPT Atlas, Perplexity Comet and Gemini's in-Chrome assistant are all live products, while every published bot document from those same operators describes the autonomous crawler instead. That gap is the whole reason for this protocol: measure the behaviour before it gets assumed and repeated as folklore, the failure mode catalogued in where AI SEO statistics come from. If the vocabulary is unfamiliar, start with the AI search glossary and what GEO is.
What is already known, and what is not
Each of these products can navigate, summarize, and take actions on a live web page at a user's direction, distinct from a chatbot's independent web search tool.
Whether they render JavaScript, respect robots.txt for the specific fetch they perform, or leave a distinguishable trace in server logs, hasn't been rigorously documented in public. The nearest thing is a 30-day single-site log study, and it does not separate the four steps below.
| Product | Operator | Interaction model |
|---|---|---|
| ChatGPT Atlas | OpenAI | Dedicated agentic browser |
| Comet | Perplexity | Dedicated agentic browser |
| Gemini in Chrome | In-browser assistant mode |
Each operator publishes bot documentation for its autonomous crawlers — OpenAI, Perplexity, Anthropic and Google — and none of it describes the agentic-browser fetch path. The AI Bot Registry collects what is documented; AI crawler statistics covers the volumes.
Why is this not the same question as AI crawling?
A retrieval or training bot's incentives favour speed and scale: fetch as many pages as cheaply as possible, generally without executing JavaScript, because running a browser engine across billions of pages does not pencil out. That cost is measured — Google needs roughly nine times longer to crawl JavaScript than HTML, with a median rendering delay of several seconds.
An agentic browser has the opposite incentive structure. It renders one page, at one moment, for one user actively waiting. That resource profile is closer to a real browser tab — the page-weight and interaction budget the Web Almanac performance chapter and the Core Web Vitals thresholds describe — than to a batch crawler. So the protocol does not assume agentic browsers behave like crawlers, even when one company runs both. Anthropic, for instance, documents an answer-time web search tool entirely separately from its training crawler.
| Dimension | Autonomous crawler | Agentic browser |
|---|---|---|
| Trigger | A schedule or a queue | One user, one request, right now |
| Scale | Billions of pages | One page at a time |
| JavaScript | Usually skipped for cost reasons | Plausibly executed — untested |
| Interaction | None | Possible: clicks, tabs, scrolling |
| robots.txt | Documented per bot | Undocumented for a user-directed fetch |
| Log signature | Declared user agent | Unknown, possibly indistinguishable from a human |
What happens when an agent opens a page?
The dependent variables come from separating four steps that public commentary collapses into one question: "does it read my page?" That question cannot be answered usefully. The four steps fail in different ways, and each has a different fix.
- 1 Fetch A request with a user agent and headers. Failure mode: the site cannot tell this is not a person.
- 2 Render Raw HTML parsed as text, or a full browser engine running the scripts. Failure mode: client-side content is invisible.
- 3 Interact Clicking a tab, expanding an accordion, scrolling for lazy content. Failure mode: content behind a click is never reached.
- 4 Extract Deciding what answers the user. Failure mode: the answer sits below three screens of preamble.
Open question Whether these products identify themselves consistently at step one is not documented publicly. Step three — interaction — is the step that separates an agent from a crawler, and the least studied of the four.
Method: sample, control, variables and schedule
Sample. One instrumented test site under our control, one page per test condition, three products: Atlas, Comet and Gemini in Chrome. Each product is run repeatedly across several days rather than once, because a single run is a snapshot rather than a behaviour.
Control. The same page visited by an ordinary human-driven browser session, logged identically, establishes what a normal visit looks like in our own logs. Without that baseline there is nothing to call an agentic fetch unusual against. A second reference point is the classic-crawler baseline from the JavaScript rendering experiment on the same markup.
Independent variable. The client: human browser, classic retrieval bot, or each agentic browser in turn, with the prompt held fixed and written down.
Dependent variables. Which of three planted strings the client can report back — one in the static HTML, one revealed only after JavaScript runs, one revealed only after a button is clicked — plus the server-side record: full user agent, headers, IP, timestamp, path, referer, and whether robots.txt was requested at all.
Analysis. Per product, never averaged across products. Each run is scored on the highest of the four steps it demonstrably reached, with counts reported as counts, not rates, until the sample is large enough to support a rate. Where server evidence and the agent's own narration disagree, the log wins and the disagreement is recorded as its own observation.
Schedule. Instrumentation first: the test page, the server-side logging and the human control session. Then a first collection window in which every product is run against every condition inside a short span of days, so a mid-window product release cannot be mistaken for a difference between products. The window is deliberately narrow for that reason, which is also why the sample stays small.
The panel then repeats on a fixed cadence rather than running once. A single reading of a product that ships weekly has a short shelf life, and a number with no expiry date circulates forever. Each repeat re-runs the identical page and prompts, so the comparison between windows — not any single window — is where the information lives.
The procedure follows the AI visibility measurement standard, is recorded under the dataset strategy, and is registered on the studies index alongside the rest of the AI Citation Index roadmap. The three-string design is the one used in zero to cited, extended by one level to cover interaction.
Pre-registered hypotheses
| # | Hypothesis | Predicted outcome |
|---|---|---|
| AB1 | Agentic browsers render JavaScript on the pages they visit, unlike autonomous retrieval bots | Expect to hold |
| AB2 | Agentic browsers respect robots.txt directives less consistently than autonomous crawlers do, given their user-directed rather than bulk-fetch nature | Expect to hold |
| AB3 | Agent self-reports of what they did on a page diverge from the server-side evidence in at least some runs | Expect to hold |
ChatGPT Atlas, Perplexity Comet, and Gemini in Chrome are already shipping. What they actually fetch from a page — and whether it behaves like a bot or like a human browser — has barely been studied.
Share on XWhat could make a result look real when it is not?
Several things can produce a pattern that looks like agent behaviour and is not. Naming them in advance is part of the method.
Caching and prefetch. A page the agent already holds may never be re-fetched, so an absent log line is not proof it ignored you. Conversely, some fetches happen before a user commits to a page; those look identical to a real read, but nobody saw the result.
Shared infrastructure. A request may come from a general fetching service used by several features, so attributing it to the agentic browser specifically may simply be wrong.
CDN and edge behaviour. Your edge may serve a cached variant, strip a header, or challenge the request — in which case origin logs tell a story about your CDN, not the agent. Blanket AI-blocking rules are easy to apply by accident, as large-scale robots.txt analysis and a dedicated blocking report both show. Our own count is the robots.txt AI-blocking census, and GPTBot vs OAI-SearchBot covers the distinction people most often get wrong when writing one.
Version drift and prompt sensitivity. A product may update mid-window, so a change in the data may be a release rather than a finding. And how the agent is asked to visit changes what it does: "summarize this page" and "find the hidden detail" are different experiments, both valid, never to be mixed.
The test page itself. A page built purely for testing has no inbound links, no history and unusual markup, and an agent may treat it differently from a real page. Every observation therefore records the date, the visible product version, and the serving path.
Null results we would publish
- No distinguishable trace. If agentic browser visits are indistinguishable from ordinary human sessions in server logs, that is a significant, publishable result.
- No interaction at all. If no product clicks the button, the interaction hypothesis fails and the "agents behave like humans" framing weakens.
- No difference from the crawler baseline. If agentic fetches match ordinary retrieval bots, the case for treating this as a separate surface collapses.
- Unstable results. If repeated runs disagree with each other, we report the instability rather than picking the run that reads best.
Refusals count too: if an agent declines to visit, or answers something unrelated, that is recorded and filed in the null results registry rather than dropped.
What would a result actually mean for a site?
The following example is invented to illustrate the mechanism. It is not a case study, and no client or measured result is being described. Take a hardware product page: the price sits in the initial HTML, the specifications sit inside a tab rendered only after a click, the warranty terms sit in an accordion. An agent asked to compare that product against two rivals can read the price and may never reach the specifications, so the answer the user receives is thin and possibly wrong by omission.
Move the specifications into the initial HTML as a plain table and nothing changes for the human; everything changes for a client that does not click. That interaction cost is not hypothetical for people either — it is why interaction latency became a Core Web Vital.
Hypothesis The prediction is that the second version is legible to more agents than the first. That prediction has not been tested here, and this page will not claim it has.
The implied fix is not detecting the agent, and not serving it a special page. It is making one page that works for a client with fewer capabilities — the same move that survived the AJAX era, and the principle underneath the technical GEO audit.
Nothing here is a ranking factor. This measures fetch and interaction behaviour. Whether an agent chooses to cite you is a separate question, covered by the cross-platform concordance protocol. Adjacent structured interfaces — MCP servers and WebMCP, placed in context by a survey of the 2026 agent-protocol stack — would change the question from "what can it render" to "what do you expose", which is why the baseline is worth taking now.
Reproduction: running this yourself
You do not need a lab. You need one page, a server log, and patience. The protocol is deliberately cheap enough to replicate.
Build the page with those three strings. Log the full user agent, IP, timestamp, path and referer at the server rather than in client-side analytics. Visit once with a normal browser to establish your control.
Then direct each agent to the page one at a time, using the same written-down prompt, noting the date and any visible product version, and ask which strings it can see. Repeat across several days. Published user-agent inventories — this crawler reference, plus OpenAI's and Perplexity's own docs — are the right starting grep list, even though none of them names an agentic-browser agent explicitly. The free tools here cover the first checks, and Claude Code for SEO covers running the whole loop on a schedule.
Raw data and what gets published
When the first collection window closes, the release is raw material rather than a summary. That means the test page source, the exact prompts used, the per-run server log lines with timestamps and user agents, the agent transcripts, and a per-product table of which of the four steps each run reached. Runs that failed or were discarded are listed with the reason, so the denominator is visible. Nothing exists yet — there is no dataset to download at this stage, and this section describes what will be published, not what has been.
Open questions
- Does a robots.txt disallow apply to a fetch a user explicitly requested? Different products may answer this differently, and none has published a clear rule.
- Do agents follow internal links, or stop at the page they were given?
- Does an agent see the same page a logged-out human sees, or does it inherit the user's session?
- How does an agent handle a consent banner, a paywall, or an age gate?
- Is there any relationship between what an agent fetches and what the same company's chatbot later cites? We have no evidence either way.
Limitations
- Sample. One test site, three products, a handful of runs per product. Counts at that scale cannot support a percentage, and this study will not report one.
- Engine coverage. Atlas, Comet and Gemini in Chrome only. Other assistants' browsing modes, enterprise agents and headless automation frameworks are outside the panel, and results will not be extrapolated to them — behaviour does not transfer between products, even between two products from the same operator.
- Page selection. A purpose-built test page with no inbound links, no history and deliberately unusual markup is not a representative page. Behaviour observed on it may not match behaviour on a real, established URL.
- Geography and session. All runs originate from a single region on a single network, from fresh accounts. Personalisation, locale routing and account history are uncontrolled and may change what an agent does.
- Measurement. Server logs record a fetch, not an intention. Caching, prefetch, shared fetching infrastructure and CDN behaviour can each hide or fabricate a request, and an agent's own narration is never treated as evidence.
- Reproducibility. These products ship changes weekly, so an identical re-run months later may legitimately produce different results. Every observation is dated and versioned for that reason; a mid-window product release is a confound we can note but not remove.
- Generalisation. Volume is unmeasured — no operator breaks out agentic browser traffic, so nothing here can be weighted by how much of it exists. The figure sets in AI SEO statistics and the state of AI search have the same gap.
To find out what your own pages expose to a client that does not click: plant the three test strings described above and read your server logs, then work through the technical GEO audit for the rest of the surface. To see which bots you are already allowing or blocking: check the identifiers in the AI bot user-agent registry.
Results, the test page source and the raw log files go out to the newsletter when the first collection window closes — that is the way to get notified when this reports. If you want to propose a test condition, offer a site to instrument, or point out a flaw in the design before it runs, the about page explains how to get in touch. There is no sign-up form for participation; a message is enough.
Namdev, R. (2026). What Agentic Browsers Actually Fetch (v1). Retrieved from https://ritiknamdev.com/blog/agentic-browsers-what-they-fetch Published under CC BY 4.0 — reuse freely with attribution.
Complements the JavaScript rendering experiment and the AI Bot Registry, which covers autonomous crawlers rather than user-driven agentic browsing.