Multiple "X% of websites block AI crawlers" figures are already in circulation. We checked several and none disclosed a reproducible sampling frame — which domains, how many, from when. Rather than add another unsourced number, this page pre-registers a census we can actually run and verify.
No data has been collected yet. Nothing on this page is a result. The sampling frame, bot list and analysis plan below are committed in advance. The census has not been run.
- Major AI crawlers publish documented user-agent strings, so blocking is directly observable in a public
robots.txtfile. - Several "X% of sites block AI crawlers" figures circulate publicly, and at least two of them are traced on this page to sources with no disclosed sample.
- Blocking policy is per-bot: a site can block a training crawler while allowing a retrieval crawler, and the two have entirely different consequences.
- What share of any defined population actually blocks each named bot.
- Whether blocking differs by publisher category, and by how much.
- Whether the blocking rate is rising, flat or falling over time.
- Whether site owners differentiate by bot purpose, or apply one blanket rule.
What question does this census answer?
What share of the web's most significant sites block each named AI crawler, and is that share moving? Nobody can currently answer it from a source you can check. Several "X% of websites block AI crawlers" figures circulate; two of them are traced below to origins that do not disclose a reproducible sample. This page pre-registers the census meant to replace them, and fixes the method before a single file is fetched — the same discipline as the AI Citation Index.
The question is answerable, which is what makes the absence of an answer notable. Major operators publish
their user-agent tokens, and robots.txt is a public file at a fixed location on every
domain. The measurement is a fetch and a parse. What has been missing is anyone willing to publish the
sampling frame and the parser rules alongside the percentage.
Where do the circulating blocking figures come from?
This part is finished work. Applying the tracing method from the provenance audit to two figures currently in circulation, both dead-end before reaching a checkable sample:
| Claim | Stated sample? | Stated date? | Grade |
|---|---|---|---|
| Coverage citing a rising share of top sites blocking at least one AI crawler | Not disclosed at the level needed to reproduce | Approximate only | Partial |
| Coverage citing a specific crawl-to-refer-derived blocking estimate | Not disclosed | Approximate only | Partial |
Neither is singled out as unusually weak; this is the category norm, which is the point. The better-documented public attempts are BuzzStream's news-site study, Paul Calvano's top-1000 robots.txt analysis and this blocking report. Each states more of its method than most; none states enough to reproduce end to end. The grading scheme behind those Partial marks is described in the measurement standard.
A meaningful and rising share of large sites block at least one major AI crawler. Consistent with the publisher-hostility narrative around crawl-to-referral economics, but the specific percentage in circulation could not be traced to a disclosed sample.
Blocking correlates with publisher category. News and media sites are widely assumed to block more aggressively than e-commerce or SaaS. Plausible; we found no cross-category breakdown with a stated method.
Method: how the census will be run
- 01 Sample domains A published top-N list, stratified by category
- 02 Fetch robots.txt Direct HTTP fetch, no cache, dated
- 03 Parse directives Per-bot Allow/Disallow, structured
- 04 Classify Blocked / allowed / no rule, per crawler
- 05 Publish raw CSV of every domain × bot × rule
Every step is mechanical and repeatable by anyone with the domain list. Fetches are direct HTTP requests to each origin, uncached and timestamped, so the result reflects the declared policy at a stated moment rather than a CDN's view of it. Rate limits are respected: a single lightweight request per domain, to a file that exists specifically to be read by automated clients.
Sample: which domains, and why that bounds the answer
A published, reproducible top-N domain ranking, stratified by category — news, e-commerce, SaaS, other. The exact N and the exact ranking source are named at collection time and published with the results, so the sample is checkable rather than asserted. What it is not is a random sample of the web. Large established properties are the group most likely to have any deliberate policy at all; the long tail almost certainly behaves differently and this design does not reach it.
Geographic coverage is bounded the same way. Top-domain rankings skew toward English-language and Western properties, so whether sites in markets with different AI-training regulation block differently is a question this edition cannot answer. It is registered as a candidate expansion, not solved.
domain: example-news.com category: news robots_txt_exists: true last_modified: 2027-01-09 GPTBot: disallow OAI-SearchBot: allow ClaudeBot: disallow Claude-SearchBot: allow PerplexityBot: no_rule collected_at: 2027-01-14T00:00:00Z
An illustrative row with invented values, shown to make the dataset shape concrete. Thousands of rows in this shape support cross-tabulation by category, bot type and presence of any rule at all — none of which a single aggregate percentage can support.
Variables: what gets recorded per domain and bot
Why per-bot classification, not one blocking flag
A single "X% block AI crawlers" figure conceals the distinction that matters most here. Training bots and retrieval bots are separate decisions with separate consequences — the choice set out in GPTBot versus OAI-SearchBot. Operators publish their own tokens: OpenAI, Anthropic, Perplexity and Google. A domain blocking GPTBot while allowing OAI-SearchBot has made a coherent choice, and it changes what each engine can quote. That is visible only if the census records per-bot outcomes.
Nor do rules transfer between crawlers. Names are per-product, so blocking one operator has no effect on another. Compliance is a per-operator property: one crawler honouring the file says nothing about the next. And a wildcard group is overridden entirely for any bot with its own named group — the misconfiguration site owners most often miss.
Pre-registered hypotheses
| # | Hypothesis | Predicted outcome |
|---|---|---|
| C1 | Training crawlers (GPTBot, ClaudeBot) are blocked more often than retrieval crawlers (OAI-SearchBot, PerplexityBot) on the same domains | Expect to hold |
| C2 | News and media domains block at a higher rate than e-commerce domains | Expect to hold |
| C3 | Blocking rate has increased since the prior year across the same domain set, re-checked | Expect to hold |
| C4 | A large majority of domains apply a uniform "block all AI" policy rather than differentiating by bot category | Expect not to hold |
C4 predicts against the intuitive assumption. Our expectation, unverified until the census runs, is that most site owners either take no action or apply a blanket rule copied from a CMS template, rather than the differentiated policy the registry recommends.
Multiple reports quote a specific percentage of sites blocking AI crawlers. None we checked disclosed a reproducible sample. So instead of repeating one, we're pre-registering a census that will.
Share on XAnalysis plan: four states, not two
| Recorded state | What the file says | What it means for citations |
|---|---|---|
| Blocked | A matching group disallows the whole site | Retrieval crawlers cannot read your current words at all |
| Blocked on some paths | A directory or pattern is disallowed, the rest is not | Partial. Counting this as "blocks AI" is how headline percentages get inflated |
| Allowed | A matching group explicitly permits access | No barrier. Not a guarantee of anything being crawled |
| No rule | No group matches this bot's token | Access by default. Absence is not a decision, and we count it without interpreting it |
Every published blocking percentage we have seen collapses these four into a binary, and the collapse is unrecoverable from the summary alone. Headline rates are therefore reported per bot and per state, never as one blended "blocks AI" figure, and always with the raw matrix beside them, the publication standard used across every study here. Where a file contains conflicting rules, the resolution rule is published in advance and the raw file body stored alongside the classification, so a reader who disagrees can reclassify from source.
Three comparisons carry the hypotheses, and each is specified now so the test cannot be chosen after seeing the data. C1 compares blocked rates for training versus retrieval crawlers within the same domain, paired rather than across separate populations, which removes the differences between sites from the comparison entirely. C2 compares blocked rates between category strata, reported as a difference in proportions with a confidence interval rather than a bare pair of percentages.
C3 compares the same domains at two collection dates, counting only domains present and fetchable in both runs, so a change in the ranking cannot masquerade as a change in policy. Domains that fail to return a parseable file are reported as their own category and excluded from denominators, never silently folded into "no rule".
Which parsing traps could corrupt the count?
Most of the difficulty is not fetching files but deciding what a file means. Each trap below produces a plausible-looking percentage, which is why we distrust percentages published without a parser description — including, in advance, our own.
| Trap | Direction of the error | Our handling |
|---|---|---|
| Token, not exact-string, matching | Either way, depending on the file | Match on published tokens, record the raw group |
| Case sensitivity | Under-counts blocks | Case-insensitive on agents, case-sensitive on paths |
| Group precedence | Over-counts blocks | Most-specific group only, never a union |
| The empty disallow | Over-counts blocks | Treated as allow-all, as the spec intends |
| Partial-path blocks | Over-counts blocks | Recorded as its own state, with the path |
| HTML served at the robots.txt location | Under-counts blocks | Response type recorded separately from directives |
| Redirects and subdomains | Both, unpredictably | Each host classified on its own file |
An adjacent file has the same problem in purer form. llms.txt is trivially countable and largely inert: a study across 300k domains found no clear citation effect, and our own reading is in does llms.txt work. Counting adoption is easy; counting consequence is the part this census deliberately does not claim.
What robots.txt does not do
Fact The mechanism is voluntary. A well-behaved crawler fetches the file, finds the group matching its user-agent, and respects it. A badly-behaved one ignores it and the site sees no error. The best-documented failure is Cloudflare catching undeclared crawlers evading no-crawl directives. The file also governs fetching, not use: a page fetched last year is not un-fetched by a rule added today, and a page syndicated elsewhere is not covered at all.
So this census measures declared policy, not enforced outcome. Those are different quantities and they can diverge widely. Request-volume evidence from network operators answers the other half of the question, and the crawler statistics page collects it.
Hypothesis Our working view is that this makes blocking a weaker instrument than either side of the debate treats it as — a signal and a friction, not a wall. We have no measurement establishing that and are not claiming one.
How should the eventual number be read?
Four misreadings are predictable enough to rule out now. "The web is closing" — a top-domain sample is not the web. "Blocking works" — the census measures declarations and cannot measure compliance; conflating them is the most damaging error available. "Blocking is the consensus" — a majority is not a consensus and a copied template is not a decision, and we cannot separate the deliberate from the inherited. "The trend will continue" — two points make a line, not a trend, so any chart will show the points rather than a fitted curve.
Nor does the result tell any individual site what to do. If most of the sample turns out to block, that is a popularity argument, not a reasoning one. Suppose instead C1 and C4 both hold: training bots blocked more than retrieval bots, but most sites still applying an undifferentiated policy. That would mean a large share of publishers have citation-driving retrieval bots blocked alongside training crawlers they meant to restrict — an easily fixable mistake, and a cheaper win than most technical SEO work.
What gets published if the hypotheses fail
If C1 fails, and training crawlers are not blocked more than retrieval crawlers, the entire framing of a differentiated publisher strategy is wrong, and that becomes the headline. If C2 fails, a widely repeated assumption about publisher behaviour loses its basis. If C3 fails, and blocking is flat or falling, that contradicts the prevailing narrative about rising resistance, and we publish the flat line. If too many domains return unparseable files, we publish the failure rate and no headline percentage at all.
Each of these is registered in advance in the null results registry, because a prediction that can only be confirmed is not a prediction.
Schedule
- v0Now
Design and pre-registration
This page. Method, hypotheses, parser rules and publication commitments, all fixed before any file is fetched.
- v1Q1 2027
First collection
Stratified top-N domain list, per-bot classification, raw CSV published alongside the summary.
- v2Semi-annual
Re-run against the same domain list
Year-over-year change becomes comparable rather than an artefact of a different sample.
- v3Later
Non-English and long-tail expansion
A candidate expansion, registered rather than promised. The initial design cannot answer it well.
First collection is targeted alongside Index v1 in Q1 2027, then re-run semi-annually against the same domain list so year-over-year change is comparable rather than an artefact of a different sample. robots.txt policy moves slowly; a monthly re-check would mostly measure noise. The panel is fixed deliberately: re-running against a freshly generated top-N list each time would confound genuine policy change with churn in the ranking itself, which is the single most common way a year-over-year crawler statistic becomes meaningless.
Raw data and reproduction instructions
The published artefact is the CSV of domain × bot × rule, with the stored file body and the collection timestamp for each row, released at the same moment as the summary. A percentage cannot be audited; a file of domains and rules can.
You can run a one-domain version in fifteen minutes, and it is worth doing before forming an opinion about anyone's numbers:
- 1 Fetch the raw file Your own robots.txt, directly. Confirm plain text and a success status. Read the bytes, not a rendered browser view.
- 2 List every user-agent group In order. Note which you added deliberately and which you inherited from a template.
- 3 Pick the one group that applies Per crawler. Most-specific match wins and groups do not stack, so only that group takes effect.
- 4 Record four states, not two Blocked entirely, blocked on some paths, allowed, or no rule — per bot.
- 5 Repeat per subdomain, then check logs Each host has its own file and its own answer. Server logs show whether a rule you never intended is blocking traffic you wanted.
Date the result and re-check after any infrastructure change. CDN configurations and CMS upgrades rewrite this file more often than people expect. The mechanical parts are covered by the free tools here.
Open questions this census cannot close
Open question Do blocked crawlers actually stop fetching? That needs server logs from many sites, not robots.txt files. A different study.
Open question Does blocking change how often a site appears in AI answers? Confounded by everything. Sites that block differ from sites that do not in many ways beyond the block, and the engines disagree with each other about sources in the first place — see the concordance study.
Open question How many rules are written by a person versus generated by a platform? Unanswerable from the file alone, and the answer would reframe the whole dataset.
Limitations
- Sample. A stratified top-N list describes large, established sites. It is not a random sample of the web and says nothing about the long tail, where most domains live.
- Geography and language. Published top-domain rankings skew English-language and US/EU. Markets with different AI-training regulation are effectively unmeasured in this edition.
- Crawler coverage. Only bots with published, stable user-agent tokens can be classified. Undeclared or rotating crawlers are invisible to this method by construction.
- Measurement. The census reads declared policy, not enforced outcome. It cannot tell you whether a disallowed bot complied.
- Confounder we cannot remove. Deliberate rules and inherited CMS-template rules are indistinguishable from the file alone, so a "block" may encode a decision or a default. The Googlebot control separates broad blocking from AI-specific blocking, but not intent from inheritance.
- Reproducibility window. A domain's robots.txt can change between the collection date and the day you read the result, so every figure carries its timestamp and the stored file body.
- Generalisation. Findings describe this domain list on this date. They do not license claims about "the web", about crawler behaviour, or about the citation consequences of blocking.
To find out whether your own robots.txt blocks the retrieval bots you want: check your file against the identifiers in the AI bot user-agent registry, which separates training crawlers from retrieval fetchers. To generate the companion file the same crawlers read: use the llms.txt generator.
The first results and the full raw CSV go out to the newsletter when the collection window closes. If you want to challenge the parser rules or the sampling frame before it runs — the most useful time to do it — the about page explains how to get in touch.
Namdev, R. (2026). The robots.txt AI-Blocking Census (v0). Retrieved from https://ritiknamdev.com/blog/robots-txt-ai-blocking-census Published under CC BY 4.0 — reuse freely with attribution.
This follows the same discipline as the AI Citation Index — publish the method before the number. See the AI Bot Registry for the crawlers this census will check for, GPTBot vs OAI-SearchBot for the decision each row encodes, and AI crawler statistics for the volume side.