Original research · Pre-registered experiment

Do AI crawlers render JavaScript?

Only Google documents that its crawler renders pages. OpenAI, Perplexity and Anthropic document identity and opt-out, not rendering. This page registers the controlled experiment that would settle it — and catalogues what is documented today.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — test site build stage ·14 min read ·Last verified September 2026
The short version

If AI retrieval bots do not execute client-side JavaScript, any page content that only appears after JS runs — the norm in single-page applications — is effectively invisible to them, however good the content is. Today only Google documents that its crawler renders pages; OpenAI, Perplexity and Anthropic document bot identity and opt-out and say nothing about rendering. Nobody has tested it publicly with a controlled experiment. This page registers that experiment and, meanwhile, lays out what is genuinely documented per bot.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. Test pages have not been built or published, so no bot has been observed against them.

What is known
  • Google documents that Googlebot renders pages using a recent version of Chrome.
  • OpenAI, Perplexity and Anthropic document bot identity and opt-out, but not rendering behaviour.
  • Rendering behaviour differs by bot — it is not one answer for all AI crawlers.
What is not yet known
  • Whether documented behaviour matches observed behaviour under a controlled test.
  • Which specific retrieval bots execute JavaScript, and to what depth.
  • Whether client-rendered content is retrievable in practice, not just fetchable.

Research question: do AI crawlers execute JavaScript?

The question: when an AI retrieval bot fetches a page whose main text is injected by client-side JavaScript, does that text ever reach the engine's answer? Every other content-tactic question on this site — llms.txt, author signals, brand mentions — is optimization at the margin. This one is binary. If the answer is no, an entire category of modern web architecture is structurally excluded from AI citation regardless of content quality. That sits upstream of everything in the technical GEO audit, and outranks anything on the GEO tactic evidence scoreboard.

Why it is still open: the operators' own documentation answers a different question. It covers who the bots are and how to block them, not what their fetch layer executes. And the one adjacent public test — five major AI systems extracting only visible HTML on live fetch — tests structured-data extraction, not script execution. That is the finding behind our schema markup study, and it is consistent with Ahrefs' correlational look at schema and AI citations. Open question Nobody has directly, publicly tested the rendering question with a controlled experiment.

What does each operator actually document today?

Deferring entirely to a future study would be unhelpful when a partial answer already exists in public documentation. Below is every crawler whose operator documentation is cited on this page, with what that documentation does and does not say about rendering. Nothing here is inferred from behaviour: a cell reads "documented" only where the operator states it, and "not documented" everywhere else. Undocumented is not the same as "does not render" — it means the public record is silent.

Current best evidence per bot, from operator documentation only. No observation from this study exists yet.
BotOperator's own documentation saysRenderingProvenance
Googlebot Google's crawler overview states Googlebot renders pages using a recent version of Chrome; AI features guidance confirms it feeds Google's AI surfaces. Documented — renders Traceable
GPTBot OpenAI's bots page documents it as the training crawler, with user-agent, IP ranges and robots.txt opt-out. Rendering is not mentioned. Not documented Open question
OAI-SearchBot Documented in the same OpenAI bots reference as the search-surfacing fetcher, separate from GPTBot. Rendering is not mentioned. Not documented Open question
ChatGPT-User Documented by OpenAI as the user-triggered fetch, again by identity and opt-out only. Not documented Open question
PerplexityBot Perplexity's crawler reference lists its crawlers, their user-agents and how to control them. Rendering is not addressed. Not documented Open question
ClaudeBot Anthropic's crawler notice covers what it crawls and how to block it. Rendering is not addressed. Not documented Open question
Claude's web search fetch The web search tool documentation describes an answer-time fetch, without stating what the fetch layer executes. Not documented Open question

One documented yes, six silences. That asymmetry is the entire case for running the experiment, and it is also why per-bot reporting is non-negotiable: a blended "AI crawlers do/don't render JavaScript" claim would be an average over one fact and six unknowns. All current agent strings are tracked in the AI bot user-agent registry.

What counts as client-only content?

Client-side rendering (CSR) means the server sends a mostly-empty HTML shell plus a JavaScript bundle, and the browser builds the page by executing that bundle — the classic single-page application. Server-side rendering (SSR) means the server sends fully-formed HTML that JavaScript then hydrates. Static generation, this site's approach, bakes the HTML at build time, and is the pattern most consistent with what Onely describes as LLM-friendly content. A bot that does not execute JavaScript sees full content under SSR or static generation, and an empty shell under pure CSR.

What did Googlebot's own JS journey cost?

This is a rerun, for AI bots, of a debate classic SEO already lived through. For years after JavaScript frameworks became common, "make sure your content is in the initial HTML" was standard advice, because Googlebot was not reliably executing it. Google eventually built the rendering pipeline it now documents — and even that introduced a two-wave process, where the initial crawl indexes raw HTML and a delayed pass picks up JavaScript-dependent content. Onely measured both halves of the gap.

9×

more time Google needed to crawl JavaScript than to crawl equivalent HTML, in Onely's measurement.

Onely
~5s

median delay Onely measured between Google's raw crawl and its rendering pass — seconds, not the weeks folklore claimed.

Onely

That cost gap is why the two kinds of bot may diverge. An index-building crawler works under a documented crawl budget and can afford to render tomorrow; nobody is waiting. A live retrieval bot is fetching while a user watches a spinner — the pattern Digital Applied observed in its 30-day agentic-crawler log study. Script execution is the slow part of a page load, which is why INP and the rest of the Core Web Vitals thresholds exist at all.

Hypothesis This economic asymmetry is our reason for predicting that live retrieval bots do not render while index-building crawlers do. It is a reasoned prior, not a finding — and if the result matches it, a reader should be more suspicious, not less.

Method: how the experiment is built

Test-site pipeline
  1. 01 Build a test site Instrumented, controlled content
  2. 02 Vary rendering Server-rendered vs. client-only per page
  3. 03 Log every fetch Per bot, full request/response
  4. 04 Check citation Does content only in JS get cited?
  5. 05 Publish per-bot results Not a blended average

Sample. A small, purpose-built test site with matched page pairs: identical content, one version rendered server-side and visible in raw HTML, one requiring client-side JavaScript execution to appear. Both are indistinguishable to a human visitor.

Control. The server-rendered twin is the control for every client-only page. A bot that fetches and cites the control but never the twin has demonstrated the difference; a bot that cites neither has demonstrated nothing, which is why the pairing matters more than the sample size.

Variables. Rendering mode is the only intended difference. Measured outcomes are: the fetch itself (from server logs, per user-agent), and whether the tracer content ever surfaces in an engine's answer. Each page carries a unique tracer sentence — the same technique the Citation Index uses for sampling — so a citation can be traced unambiguously back to one variant. If an engine cites the tracer from the JS-only page, the bot behind it executed the JavaScript. If it never does while citing the server-rendered equivalent, it did not.

Hypothesis: what we predict before running it

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
JS1Retrieval bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot) do not execute JavaScript and never cite JS-only contentExpect to hold
JS2Googlebot (feeding AI Overviews and AI Mode) does render JavaScript, unlike the non-Google retrieval botsExpect to hold

Only Google documents that its crawler renders pages. OpenAI, Perplexity and Anthropic document identity and opt-out, not rendering. One documented yes, six silences — and no controlled public test. We're registering the one that would settle it.

Share on X

Control: confounds that could fake either answer

A clean-looking result can still be wrong. Each of these is a way the measurement could mislead, and each has to be excluded before any conclusion is drawn.

The bot never fetched the page. Absence from answers proves nothing about rendering if the request never happened. Server logs must confirm the fetch first.

The page was fetched but judged irrelevant. Non-citation is weak evidence. The tracer fact has to be the obvious answer to the query, or the test measures relevance rather than rendering. Queries that Wikipedia already answers well will never reach a test page at all.

Caching. An engine may answer from a copy taken before the JS content existed. The pages have to sit long enough to be recrawled — the interval measured directly in the crawl-to-citation latency study.

The answer came from somewhere else. If the tracer text appears anywhere else on the web, the citation may not be ours. Tracer uniqueness is a design requirement, not a nicety.

The JavaScript itself failed. If the injection script errors in the bot's environment for reasons unrelated to rendering policy, a negative is an artefact. The control is the same page in a real browser.

Analysis: how a result would be read, and misread

Results are reported per bot and never blended, the same discipline applied in the cross-platform concordance work. Rendering capability differs per operator, and so do timeout budgets, cache behaviour and fallback paths through a conventional search index. Which of an operator's several crawlers fetched the page matters too, as Cloudflare's split of training versus retrieval traffic shows and the GPTBot versus OAI-SearchBot split illustrates. A result about one bot transfers to another only by assumption.

Three misreadings are worth pre-empting. "AI can't read JavaScript sites" is too broad. The claim under test concerns specific bots, at a specific date, for content that exists only after client execution. "So SSR is a ranking factor" confuses eligibility with preference — a collapse the statistics-provenance audit keeps finding. Nothing in this design can say whether server rendering improves position among pages already eligible. "The result is permanent" ignores that crawler behaviour is software: any finding carries a date and has to be re-run.

Interpretation: null results we would publish

A study that can only report one kind of outcome is not a study. Each of these lands in the null results registry alongside the rest of our studies.

  • Every tested bot renders JavaScript fine. That contradicts our stated prior, and we would say plainly that our reasoning about crawler economics was wrong.
  • Rendering makes no difference to citation. Both variants cited at similar rates — a genuinely useful and completely boring finding.
  • Too few fetches to conclude anything. We report the fetch counts and stop, rather than reason from a handful of events.
  • Inconsistent behaviour with no clean pattern. Some renders, some not, no stable rule. The hardest outcome to write up, and still published.

Which architectures are exposed, and which are not

SituationDoes this matter?
Static site generator, content in HTML at build timeNo. Nothing here changes what you should do.
Server-rendered app with client hydrationMostly no, if the text is in the first response. Verify rather than assume.
Pure client-rendered SPA, empty shell in the responseYes. This is the exposed case.
Content behind a client-side tab or accordionDepends. If the text is in the DOM and merely hidden by CSS, it is present. If it is fetched on click, it is not.
Third-party embedded widgets carrying key factsYes. Embedded content is frequently absent from the host page's HTML — a recurring problem for YMYL pages whose disclosures live in a widget, and for commerce listings whose price and stock blocks are injected client-side.
Content loaded by infinite scrollYes, for everything past the first batch.
An MCP server or WebMCP endpointNo. Agent-facing interfaces sidestep HTML rendering entirely.

The mitigations, if you are in an exposed row, are not exotic. They are long-standing technical-SEO practices that predate AI search and overlap almost entirely with the work catalogued in the HTTP Archive's 2025 performance chapter.

Server-side rendering (SSR)Full page markup, including dynamic content, generated on the server before any client JS runs.
Static generationContent baked into HTML at build time — this site's own approach.
Dynamic rendering / prerenderingServe a fully-rendered snapshot to known bot user-agents, while humans get the interactive client-rendered version.
Pure client-side rendering with no fallbackThe configuration this experiment is designed to test for citation risk — content that exists only after JS execution, with no server-rendered equivalent.

Reproduction: check your own site in thirty minutes

You do not need this study to answer the version of the question that affects you. Our free tools include a checker that does the plain fetch for you; the manual procedure is six steps.

A per-site rendering check anyone can reproduce
  1. 1 Pick three pages that carry citable content Not the homepage. The pages you want quoted — start with the ones most likely to be cited.
  2. 2 Fetch each with a plain HTTP client No browser. Save the raw response to a file.
  3. 3 Search the file for a distinctive on-page sentence Something you can see with your own eyes in a browser. Record present or absent.
  4. 4 Repeat with an AI crawler user-agent Momentic keeps a usable list of current agent strings. Some sites serve different markup by agent, sometimes by accident — a difference here is itself a finding.
  5. 5 Check server logs for those agents Confirm the bots reach the pages at all. A rendering question is moot if nothing is fetching you.
  6. 6 Date the result and re-run after any front-end release Your rendering behaviour changes more often than the crawler's does.

Agent strings come from Momentic's crawler reference; for a sense of what normal fetch volume looks like, Cloudflare's 2025 crawler census is the best public benchmark, and blocking decisions are tracked in the robots.txt blocking census. This procedure has one obvious limit. It tells you what a bot could see, not what it did see. That is the gap the zero-to-cited log study was built to close, and the gap the controlled experiment exists to close properly.

Raw data and how to make the case internally

When the study reports, three files publish alongside it. The per-bot server log extract shows which user-agents fetched which variant and when. The per-query citation checks carry timestamps. The test site's own page source for both variants lets anyone rebuild the pair. Nothing is summarised without the underlying rows, which is what the measurement standard asks of everyone else.

If your own thirty-minute check comes back negative, the conversation with whoever controls engineering time goes better in this order.

Lead with the observable fact — "our product pages return an empty shell to any client that does not run JavaScript." It is checkable in one command and nobody can argue with it. Then state the risk conditionally, and say out loud that it is not yet publicly established for non-Google bots. Google's AI optimization guide is the only first-party statement worth quoting, and it settles nothing for anyone else.

Then scope the work honestly. Adding server rendering to a mature single-page app is not a weekend task, while a prerendered snapshot for known bots is a smaller intervention worth pricing separately. Finally, name the metric you will report afterwards. Raw-HTML content presence is honest and within your control. Citation counts are not, and promising them costs you the next request.

Limitations

  • Sample: one test site. Results may not generalise across every JS implementation pattern, and a crawler could plausibly spend more rendering budget on domains it already trusts. A single test domain cannot detect that; reproduction on a second site with different infrastructure is the fix, and we would want it run before anyone acts on a headline number.
  • Engine coverage is partial. Only bots that fetch the test site in observable numbers produce a result. If a crawler is not in the published table, we have no result for it, and we will not extrapolate from one we tested.
  • Query selection drives non-citation. The tracer must be the obvious answer to the test query. Where an incumbent source answers the query well, the test page is never reached and the null is uninformative.
  • Measurement cannot separate "chose not to render" from "tried and timed out". Both look identical from outside. Nor can a short study detect a second, slower rendering pass arriving weeks later; that would read as a false negative.
  • Confounders remain after mitigation. Caching, index fallback paths and per-operator multi-crawler setups can each carry content into an answer without the tested bot rendering anything.
  • Reproducibility depends on the tracer staying unique. If the tracer text is scraped and republished elsewhere, the citation chain breaks and that arm of the study is void.
  • Generalisation over time is limited. Bot rendering behaviour can change without notice, so any result is dated snapshot, not a permanent architectural fact.
  • Geography and infrastructure are fixed. The test site is served from one stack in one region; behaviour under different CDN, bot-management or regional routing conditions is untested.
When this runs

Per-bot results and the raw server logs go out to the newsletter when the test site has been live long enough to draw fetches. If you want to flag a hole in the design, or you run a site whose logs would strengthen it, the about page explains how to get in touch. It sits inside the AI Citation Index programme.

How to cite this
Namdev, R. (2026). Do AI crawlers render JavaScript? (v1). Retrieved from https://ritiknamdev.com/blog/do-ai-crawlers-render-javascript

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Part of the AI Citation Index programme. Relevant background in the schema RCT's live-fetch findings; see the schema study and the AI Bot Registry.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Onely — Google needs 9x more time to crawl JS than HTMLwww.onely.com/blog/google-needs-9x-more-time-to-crawl-js-than-html Onely — Google's rendering delay: the five-second myth, measuredwww.onely.com/blog/googles-rendering-delay-5-seconds Onely — What makes content LLM-friendlywww.onely.com/blog/llm-friendly-content Google Search Central — Overview of Google crawlers and fetchersdevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers Google Search Central — Managing crawl budget for large sitesdevelopers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget Google Search Central — AI features and your websitedevelopers.google.com/search/docs/appearance/ai-features Google Search Central — AI optimization guidedevelopers.google.com/search/docs/fundamentals/ai-optimization-guide Google Search Central — Core Web Vitalsdevelopers.google.com/search/docs/appearance/core-web-vitals web.dev — Defining the Core Web Vitals thresholdsweb.dev/articles/defining-core-web-vitals-thresholds web.dev — INP becomes a Core Web Vitalweb.dev/blog/inp-cwv-march-12 HTTP Archive Web Almanac 2025 — Performance chapteralmanac.httparchive.org/en/2025/performance OpenAI — Bots and crawlers documentationplatform.openai.com/docs/bots OpenAI developer docs — Bots referencedevelopers.openai.com/api/docs/bots Perplexity — Crawler documentationdocs.perplexity.ai/docs/resources/perplexity-crawlers Anthropic Support — Does Anthropic crawl data from the web?support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Anthropic — Web search tool documentationplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Cloudflare — From Googlebot to GPTBot: who is crawling your site in 2025blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Cloudflare — Crawlers, clicks and AI botsblog.cloudflare.com/crawlers-click-ai-bots-training Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots Digital Applied — 30-day agentic crawler behaviour log studywww.digitalapplied.com/blog/agentic-crawler-behavior-30-day-site-log-study Ahrefs — Schema markup and AI citationsahrefs.com/blog/schema-ai-citations
FAQ

Frequently asked questions

So do AI crawlers render JavaScript or not?
Of the crawlers documented on this page, only Google publicly states that its crawler renders pages with a recent version of Chrome. OpenAI, Perplexity and Anthropic document bot identity, IP ranges and opt-out, but do not state rendering behaviour either way. So the honest answer today is: one documented yes, and several undocumented unknowns. That gap is what the experiment is designed to close.
Why is this framed as an experiment rather than a statistics page?
Because it requires a controlled test site with known content in known rendering states. An observational statistics page cannot isolate this the way a purpose-built experiment can.
Don’t we already know AI crawlers ignore JavaScript, from the schema findings?
A related but separate test, JSON-LD extraction on live fetch, found systems extracting visible HTML only. That is suggestive, but not a direct test of whether a bot executes JavaScript to reveal client-rendered text. This experiment tests that directly.
Could a bot render JavaScript sometimes but not always?
Yes, and that is one of the outcomes we consider most likely. Rendering costs money and time. A system under load may skip it; a system with a cached page may reuse an older render. If we observe partial or inconsistent rendering, we will report the rate rather than force a binary verdict.
If a bot does render JavaScript, does that mean my single-page app is safe?
No. Rendering is only the first gate. The content still has to be retrieved, chunked, judged relevant and selected for the answer. Passing the rendering gate removes one failure mode, not the others.
Why not just ask the AI companies directly?
We read their public documentation, and this page catalogues it. But documentation describes intent, not behaviour under load, and it goes stale quietly. A measurement disagrees with a document in a way anyone else can check.
Does this question apply to Googlebot the same way?
Not exactly. Googlebot has a documented rendering pipeline that classic SEO has studied for years, so the open question there is timing rather than capability. For the newer retrieval bots, capability itself is unestablished in public. The two cases are kept separate throughout.
Can you tell the difference between a bot choosing not to render and one timing out?
Not with this design. Both look identical from outside: content that was fetched and never cited. We will state that limit in the results rather than paper over it.
Is there a quick way to check this for my own site before the study reports?
Fetch your page with a plain HTTP request, no browser, and check whether your key content appears in the raw response. If it only appears in a real browser, that content is at risk under exactly the hypothesis being tested here.
What would make you abandon or retract this study?
Too few bot fetches to conclude anything; evidence that the tracer method leaks, so a citation could arrive by a route other than the page we served; or a failure to reproduce on a second test site with different infrastructure. Each of those outcomes gets published as-is.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.