Original research · Pre-registered

The robots.txt AI-Blocking Census

How many of the web's most important sites block AI crawlers, and how that's changing — measured directly from robots.txt files rather than quoted from a report with no visible sample.

Several figures already circulate claiming a specific share of sites block AI bots. None of the versions we checked disclosed a reproducible sample. This page is the pre-registration for an independent census that will.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — design stage ·14 min read ·Last verified September 2026
The short version

Multiple "X% of websites block AI crawlers" figures are already in circulation. We checked several and none disclosed a reproducible sampling frame — which domains, how many, from when. Rather than add another unsourced number, this page pre-registers a census we can actually run and verify.

Research status Protocol — pre-registered

No data has been collected yet. Nothing on this page is a result. The sampling frame, bot list and analysis plan below are committed in advance. The census has not been run.

What is known
  • Major AI crawlers publish documented user-agent strings, so blocking is directly observable in a public robots.txt file.
  • Several "X% of sites block AI crawlers" figures circulate publicly, and at least two of them are traced on this page to sources with no disclosed sample.
  • Blocking policy is per-bot: a site can block a training crawler while allowing a retrieval crawler, and the two have entirely different consequences.
What is not yet known
  • What share of any defined population actually blocks each named bot.
  • Whether blocking differs by publisher category, and by how much.
  • Whether the blocking rate is rising, flat or falling over time.
  • Whether site owners differentiate by bot purpose, or apply one blanket rule.

What question does this census answer?

What share of the web's most significant sites block each named AI crawler, and is that share moving? Nobody can currently answer it from a source you can check. Several "X% of websites block AI crawlers" figures circulate; two of them are traced below to origins that do not disclose a reproducible sample. This page pre-registers the census meant to replace them, and fixes the method before a single file is fetched — the same discipline as the AI Citation Index.

The question is answerable, which is what makes the absence of an answer notable. Major operators publish their user-agent tokens, and robots.txt is a public file at a fixed location on every domain. The measurement is a fetch and a parse. What has been missing is anyone willing to publish the sampling frame and the parser rules alongside the percentage.

Where do the circulating blocking figures come from?

This part is finished work. Applying the tracing method from the provenance audit to two figures currently in circulation, both dead-end before reaching a checkable sample:

Provenance trace completed before this census was designed. The grades describe the sources, not our own results.
ClaimStated sample?Stated date?Grade
Coverage citing a rising share of top sites blocking at least one AI crawlerNot disclosed at the level needed to reproduceApproximate onlyPartial
Coverage citing a specific crawl-to-refer-derived blocking estimateNot disclosedApproximate onlyPartial

Neither is singled out as unusually weak; this is the category norm, which is the point. The better-documented public attempts are BuzzStream's news-site study, Paul Calvano's top-1000 robots.txt analysis and this blocking report. Each states more of its method than most; none states enough to reproduce end to end. The grading scheme behind those Partial marks is described in the measurement standard.

Hypothesis

A meaningful and rising share of large sites block at least one major AI crawler. Consistent with the publisher-hostility narrative around crawl-to-referral economics, but the specific percentage in circulation could not be traced to a disclosed sample.

Hypothesis

Blocking correlates with publisher category. News and media sites are widely assumed to block more aggressively than e-commerce or SaaS. Plausible; we found no cross-category breakdown with a stated method.

Method: how the census will be run

Collection pipeline
  1. 01 Sample domains A published top-N list, stratified by category
  2. 02 Fetch robots.txt Direct HTTP fetch, no cache, dated
  3. 03 Parse directives Per-bot Allow/Disallow, structured
  4. 04 Classify Blocked / allowed / no rule, per crawler
  5. 05 Publish raw CSV of every domain × bot × rule

Every step is mechanical and repeatable by anyone with the domain list. Fetches are direct HTTP requests to each origin, uncached and timestamped, so the result reflects the declared policy at a stated moment rather than a CDN's view of it. Rate limits are respected: a single lightweight request per domain, to a file that exists specifically to be read by automated clients.

Sample: which domains, and why that bounds the answer

A published, reproducible top-N domain ranking, stratified by category — news, e-commerce, SaaS, other. The exact N and the exact ranking source are named at collection time and published with the results, so the sample is checkable rather than asserted. What it is not is a random sample of the web. Large established properties are the group most likely to have any deliberate policy at all; the long tail almost certainly behaves differently and this design does not reach it.

Geographic coverage is bounded the same way. Top-domain rankings skew toward English-language and Western properties, so whether sites in markets with different AI-training regulation block differently is a question this edition cannot answer. It is registered as a candidate expansion, not solved.

domain: example-news.com
category: news
robots_txt_exists: true
last_modified: 2027-01-09
GPTBot: disallow
OAI-SearchBot: allow
ClaudeBot: disallow
Claude-SearchBot: allow
PerplexityBot: no_rule
collected_at: 2027-01-14T00:00:00Z

An illustrative row with invented values, shown to make the dataset shape concrete. Thousands of rows in this shape support cross-tabulation by category, bot type and presence of any rule at all — none of which a single aggregate percentage can support.

Variables: what gets recorded per domain and bot

Per domainDomain, category, whether robots.txt exists at all, HTTP response type, and last-modified date if exposed.
Per crawlerBlocked / blocked on some paths / allowed / no explicit rule, checked against every bot in the AI bot registry.
ControlGooglebot, a long-established crawler, is classified on every domain alongside the AI bots. It is the baseline: a domain blocking Googlebot too is blocking broadly, not making an AI-specific choice.
OutputRaw CSV of domain × bot × rule, plus the raw file body, plus the aggregate report.

Why per-bot classification, not one blocking flag

A single "X% block AI crawlers" figure conceals the distinction that matters most here. Training bots and retrieval bots are separate decisions with separate consequences — the choice set out in GPTBot versus OAI-SearchBot. Operators publish their own tokens: OpenAI, Anthropic, Perplexity and Google. A domain blocking GPTBot while allowing OAI-SearchBot has made a coherent choice, and it changes what each engine can quote. That is visible only if the census records per-bot outcomes.

Nor do rules transfer between crawlers. Names are per-product, so blocking one operator has no effect on another. Compliance is a per-operator property: one crawler honouring the file says nothing about the next. And a wildcard group is overridden entirely for any bot with its own named group — the misconfiguration site owners most often miss.

Pre-registered hypotheses

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
C1Training crawlers (GPTBot, ClaudeBot) are blocked more often than retrieval crawlers (OAI-SearchBot, PerplexityBot) on the same domainsExpect to hold
C2News and media domains block at a higher rate than e-commerce domainsExpect to hold
C3Blocking rate has increased since the prior year across the same domain set, re-checkedExpect to hold
C4A large majority of domains apply a uniform "block all AI" policy rather than differentiating by bot categoryExpect not to hold

C4 predicts against the intuitive assumption. Our expectation, unverified until the census runs, is that most site owners either take no action or apply a blanket rule copied from a CMS template, rather than the differentiated policy the registry recommends.

Multiple reports quote a specific percentage of sites blocking AI crawlers. None we checked disclosed a reproducible sample. So instead of repeating one, we're pre-registering a census that will.

Share on X

Analysis plan: four states, not two

Recorded stateWhat the file saysWhat it means for citations
BlockedA matching group disallows the whole siteRetrieval crawlers cannot read your current words at all
Blocked on some pathsA directory or pattern is disallowed, the rest is notPartial. Counting this as "blocks AI" is how headline percentages get inflated
AllowedA matching group explicitly permits accessNo barrier. Not a guarantee of anything being crawled
No ruleNo group matches this bot's tokenAccess by default. Absence is not a decision, and we count it without interpreting it

Every published blocking percentage we have seen collapses these four into a binary, and the collapse is unrecoverable from the summary alone. Headline rates are therefore reported per bot and per state, never as one blended "blocks AI" figure, and always with the raw matrix beside them, the publication standard used across every study here. Where a file contains conflicting rules, the resolution rule is published in advance and the raw file body stored alongside the classification, so a reader who disagrees can reclassify from source.

Three comparisons carry the hypotheses, and each is specified now so the test cannot be chosen after seeing the data. C1 compares blocked rates for training versus retrieval crawlers within the same domain, paired rather than across separate populations, which removes the differences between sites from the comparison entirely. C2 compares blocked rates between category strata, reported as a difference in proportions with a confidence interval rather than a bare pair of percentages.

C3 compares the same domains at two collection dates, counting only domains present and fetchable in both runs, so a change in the ranking cannot masquerade as a change in policy. Domains that fail to return a parseable file are reported as their own category and excluded from denominators, never silently folded into "no rule".

Which parsing traps could corrupt the count?

Most of the difficulty is not fetching files but deciding what a file means. Each trap below produces a plausible-looking percentage, which is why we distrust percentages published without a parser description — including, in advance, our own.

TrapDirection of the errorOur handling
Token, not exact-string, matchingEither way, depending on the fileMatch on published tokens, record the raw group
Case sensitivityUnder-counts blocksCase-insensitive on agents, case-sensitive on paths
Group precedenceOver-counts blocksMost-specific group only, never a union
The empty disallowOver-counts blocksTreated as allow-all, as the spec intends
Partial-path blocksOver-counts blocksRecorded as its own state, with the path
HTML served at the robots.txt locationUnder-counts blocksResponse type recorded separately from directives
Redirects and subdomainsBoth, unpredictablyEach host classified on its own file

An adjacent file has the same problem in purer form. llms.txt is trivially countable and largely inert: a study across 300k domains found no clear citation effect, and our own reading is in does llms.txt work. Counting adoption is easy; counting consequence is the part this census deliberately does not claim.

What robots.txt does not do

Fact The mechanism is voluntary. A well-behaved crawler fetches the file, finds the group matching its user-agent, and respects it. A badly-behaved one ignores it and the site sees no error. The best-documented failure is Cloudflare catching undeclared crawlers evading no-crawl directives. The file also governs fetching, not use: a page fetched last year is not un-fetched by a rule added today, and a page syndicated elsewhere is not covered at all.

So this census measures declared policy, not enforced outcome. Those are different quantities and they can diverge widely. Request-volume evidence from network operators answers the other half of the question, and the crawler statistics page collects it.

Hypothesis Our working view is that this makes blocking a weaker instrument than either side of the debate treats it as — a signal and a friction, not a wall. We have no measurement establishing that and are not claiming one.

How should the eventual number be read?

Costs of blockingNo attributed quotation, no referral traffic from citations, and no ability to see your own content represented correctly. You may still be described secondhand, less accurately.
Costs of allowingYour content can be summarised in place of a visit, crawl volume costs bandwidth, and you forfeit some negotiating position if licensing matters to you.
The asymmetry nobody controlsBoth choices are reversible in the file and irreversible in effect. Content already fetched stays fetched.
The only confident recommendationDecide deliberately, per bot, write it down, and re-verify after every deploy. This does not depend on any number the census produces.

Four misreadings are predictable enough to rule out now. "The web is closing" — a top-domain sample is not the web. "Blocking works" — the census measures declarations and cannot measure compliance; conflating them is the most damaging error available. "Blocking is the consensus" — a majority is not a consensus and a copied template is not a decision, and we cannot separate the deliberate from the inherited. "The trend will continue" — two points make a line, not a trend, so any chart will show the points rather than a fitted curve.

Nor does the result tell any individual site what to do. If most of the sample turns out to block, that is a popularity argument, not a reasoning one. Suppose instead C1 and C4 both hold: training bots blocked more than retrieval bots, but most sites still applying an undifferentiated policy. That would mean a large share of publishers have citation-driving retrieval bots blocked alongside training crawlers they meant to restrict — an easily fixable mistake, and a cheaper win than most technical SEO work.

What gets published if the hypotheses fail

If C1 fails, and training crawlers are not blocked more than retrieval crawlers, the entire framing of a differentiated publisher strategy is wrong, and that becomes the headline. If C2 fails, a widely repeated assumption about publisher behaviour loses its basis. If C3 fails, and blocking is flat or falling, that contradicts the prevailing narrative about rising resistance, and we publish the flat line. If too many domains return unparseable files, we publish the failure rate and no headline percentage at all.

Each of these is registered in advance in the null results registry, because a prediction that can only be confirmed is not a prediction.

Schedule

Planned editions, fixed in advance
  1. v0Now

    Design and pre-registration

    This page. Method, hypotheses, parser rules and publication commitments, all fixed before any file is fetched.

  2. v2Semi-annual

    Re-run against the same domain list

    Year-over-year change becomes comparable rather than an artefact of a different sample.

  3. v3Later

    Non-English and long-tail expansion

    A candidate expansion, registered rather than promised. The initial design cannot answer it well.

First collection is targeted alongside Index v1 in Q1 2027, then re-run semi-annually against the same domain list so year-over-year change is comparable rather than an artefact of a different sample. robots.txt policy moves slowly; a monthly re-check would mostly measure noise. The panel is fixed deliberately: re-running against a freshly generated top-N list each time would confound genuine policy change with churn in the ranking itself, which is the single most common way a year-over-year crawler statistic becomes meaningless.

Raw data and reproduction instructions

The published artefact is the CSV of domain × bot × rule, with the stored file body and the collection timestamp for each row, released at the same moment as the summary. A percentage cannot be audited; a file of domains and rules can.

You can run a one-domain version in fifteen minutes, and it is worth doing before forming an opinion about anyone's numbers:

Reproduce this on your own domain
  1. 1 Fetch the raw file Your own robots.txt, directly. Confirm plain text and a success status. Read the bytes, not a rendered browser view.
  2. 2 List every user-agent group In order. Note which you added deliberately and which you inherited from a template.
  3. 3 Pick the one group that applies Per crawler. Most-specific match wins and groups do not stack, so only that group takes effect.
  4. 4 Record four states, not two Blocked entirely, blocked on some paths, allowed, or no rule — per bot.
  5. 5 Repeat per subdomain, then check logs Each host has its own file and its own answer. Server logs show whether a rule you never intended is blocking traffic you wanted.

Date the result and re-check after any infrastructure change. CDN configurations and CMS upgrades rewrite this file more often than people expect. The mechanical parts are covered by the free tools here.

Open questions this census cannot close

Open question Do blocked crawlers actually stop fetching? That needs server logs from many sites, not robots.txt files. A different study.

Open question Does blocking change how often a site appears in AI answers? Confounded by everything. Sites that block differ from sites that do not in many ways beyond the block, and the engines disagree with each other about sources in the first place — see the concordance study.

Open question How many rules are written by a person versus generated by a platform? Unanswerable from the file alone, and the answer would reframe the whole dataset.

Limitations

  • Sample. A stratified top-N list describes large, established sites. It is not a random sample of the web and says nothing about the long tail, where most domains live.
  • Geography and language. Published top-domain rankings skew English-language and US/EU. Markets with different AI-training regulation are effectively unmeasured in this edition.
  • Crawler coverage. Only bots with published, stable user-agent tokens can be classified. Undeclared or rotating crawlers are invisible to this method by construction.
  • Measurement. The census reads declared policy, not enforced outcome. It cannot tell you whether a disallowed bot complied.
  • Confounder we cannot remove. Deliberate rules and inherited CMS-template rules are indistinguishable from the file alone, so a "block" may encode a decision or a default. The Googlebot control separates broad blocking from AI-specific blocking, but not intent from inheritance.
  • Reproducibility window. A domain's robots.txt can change between the collection date and the day you read the result, so every figure carries its timestamp and the stored file body.
  • Generalisation. Findings describe this domain list on this date. They do not license claims about "the web", about crawler behaviour, or about the citation consequences of blocking.
Where to go next

To find out whether your own robots.txt blocks the retrieval bots you want: check your file against the identifiers in the AI bot user-agent registry, which separates training crawlers from retrieval fetchers. To generate the companion file the same crawlers read: use the llms.txt generator.

When this runs

The first results and the full raw CSV go out to the newsletter when the collection window closes. If you want to challenge the parser rules or the sampling frame before it runs — the most useful time to do it — the about page explains how to get in touch.

How to cite this
Namdev, R. (2026). The robots.txt AI-Blocking Census (v0). Retrieved from https://ritiknamdev.com/blog/robots-txt-ai-blocking-census

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

This follows the same discipline as the AI Citation Index — publish the method before the number. See the AI Bot Registry for the crawlers this census will check for, GPTBot vs OAI-SearchBot for the decision each row encodes, and AI crawler statistics for the volume side.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

SEOmator — crawl-to-refer ratio and AI crawler data reportseomator.com/blog/crawl-to-refer-ratio-ai-crawlers-llm-bots TechnologyChecker — robots.txt AI crawler blocking reporttechnologychecker.io/blog/robots-txt-ai-crawlers-blocking-report BuzzStream — Which news sites block AI crawlerswww.buzzstream.com/blog/publishers-block-ai-study Paul Calvano — AI bots and robots.txt, top-1000 analysispaulcalvano.com/2025-08-21-ai-bots-and-robots-txt Cloudflare — 2025 Radar Year in Reviewblog.cloudflare.com/radar-2025-year-in-review Cloudflare — From Googlebot to GPTBot: who is crawling your siteblog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Cloudflare — The crawl-to-click gap for AI botsblog.cloudflare.com/crawlers-click-ai-bots-training Cloudflare — Perplexity stealth crawling investigationblog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives Search Engine Journal — Cloudflare report, Googlebot tops AI crawler trafficwww.searchenginejournal.com/cloudflare-report-googlebot-tops-ai-crawler-traffic/563303 InfoQ — Cloudflare 2025 data on AI botsinfoq.com/news/2025/12/cloudflare-2025-ai-bots OpenAI — Overview of OpenAI crawlersdevelopers.openai.com/api/docs/bots OpenAI — Platform bots documentationplatform.openai.com/docs/bots Anthropic — Does Anthropic crawl the web, and how to block itsupport.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Perplexity — Official crawler documentationdocs.perplexity.ai/docs/resources/perplexity-crawlers Am I Cited — Google-Extended: what it does and should you block itwww.amicited.com/blog/google-extended-what-it-does-should-you-block-it Google Search Central — Overview of Google crawlers and fetchersdevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers Google Search Central — Managing crawl budget for large sitesdevelopers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots Digital Applied — 30-day agentic crawler behaviour log studywww.digitalapplied.com/blog/agentic-crawler-behavior-30-day-site-log-study llms.txt — The proposed standardllmstxt.org Ahrefs — llms.txt study across 300k domainsahrefs.com/blog/llmstxt-study SEO Sherpa — 97% of llms.txt files are never readseosherpa.com/97-of-llms-txt-files-are-never-read
FAQ

Frequently asked questions

Has this census actually been run yet?
No. This page is the design and pre-registration, published before collection, for the same reason the AI Citation Index is. A census run without a published method invites exactly the cherry-picking this site exists to push back against.
Why not just quote one of the existing "X% of sites block AI crawlers" figures?
Because those figures vary by which crawlers are counted, which domain list was sampled, and when. None of the versions we found disclosed their sampling frame clearly enough to reproduce. That gap is exactly what this census is designed to close.
What domain list will be used?
A published, reproducible top-domain ranking, not an arbitrary or undisclosed list. The specific source will be named at collection time, so the sample itself is checkable.
How is this different from Cloudflare's reporting?
Cloudflare reports request volume and behavior across its own network. This census reads the declared policy, robots.txt, directly from each domain's origin, independent of any single CDN's customer base. A different, complementary measurement.
Will the census distinguish between a deliberate block and a default template setting?
Not directly from the robots.txt file alone, since a rule looks identical whether a site owner wrote it deliberately or copied it from a CMS default. This is registered as a limitation, not solved. A follow-up survey of site owners is a candidate way to address it in a later edition.
Why measure this semi-annually rather than more often?
robots.txt policy changes slowly relative to how often other AI-search metrics move. A monthly re-check would mostly measure noise. Semi-annual re-measurement against the same domain list is frequent enough to catch a real trend without over-sampling a slow-moving variable.
Could this census itself get blocked or rate-limited while collecting data?
Fetching a domain's own robots.txt file is a single, lightweight, publicly-intended request. The file exists specifically to be read by automated clients. So this is a low risk relative to a full-site crawl. But collection will still respect reasonable rate limits, to avoid any appearance of the kind of extractive behavior this site's own research criticizes elsewhere.
Does blocking a training crawler protect my content from being used for training?
Not reliably, and the census cannot tell you otherwise. robots.txt is a request, not a lock. It also only governs the crawler that reads it. Content already collected, content republished elsewhere, and content reached through a third-party dataset are all outside its reach. Treat it as a stated preference with real but partial effect.
If most sites turn out to block, should I block too?
No. That is a popularity argument, not a reasoning one. The census measures what the web does, not what you should do. Your decision depends on where your traffic comes from, whether citations are worth anything to your business, and whether you have a licensing position to protect. Those inputs are yours, not the median site owner's.
What counts as blocked when a site has conflicting rules?
This is the hardest classification problem in the census. A file can allow a path and disallow its parent, or name a bot in two groups. We will publish our resolution rule in advance and record the raw file alongside the classification, so anyone who disagrees with our rule can reclassify from the source.
Why publish the raw CSV rather than just the summary?
Because the summary is the part most likely to be wrong, and the least checkable. A percentage cannot be audited. A file of domains and rules can. Publishing the raw table is the only way a reader can disagree with our conclusion using our own data.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.