Blog / The sites that block AI crawlers are the ones with editors
The sites that block AI crawlers are the ones with editors
A robots.txt study posted to arXiv in October 2025 reports 60.0 percent of reputable news sites disallowing at least one AI crawler against 9.1 percent of misinformation sites, and an average of 15.5 named AI user agents against 0.77.
In short
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web, posted to arXiv as 2510.10315 on 11 October 2025 by Nicolas Steinacker-Olsztyn, Devashish Gosain and Ha Dao, reports that 60.0 percent of reputable news sites disallow at least one AI crawler against 9.1 percent of misinformation sites.
- The same paper reports reputable news sites naming an average of 15.5 distinct AI user agents in robots.txt against 0.77 for misinformation sites, counted across the 3,179 reputable and 493 misinformation files that were valid when fetched.
- Its longitudinal pass over six Internet Archive snapshots from September 2023 to May 2025 records reputable AI blocking rising from 23 percent to nearly 60 percent, while misinformation sites moved only from 4.6 percent to 9.2 percent.
- At the response layer the gap nearly closes for one crawler: the paper reports 24.8 percent of reputable sites and 25.6 percent of misinformation sites appearing to actively block ClaudeBot, against 18.0 percent and 10.2 percent for Anthropic-AI.
- Lantad did not run this study, holds no capture of any site in its dataset and measured none of these figures, which are reported from the paper as read on 3 August 2026.
Advice about AI crawlers is nearly always addressed to one site at a time. Decide whether you want your pages inside a model, write the robots.txt rules that say so, and the question looks settled for your domain. A study of robots.txt gatekeeping asks a different question, which is what the corpus looks like once a large population of sites has made that decision deliberately and another population has never considered it at all. The answer is a split far wider than any per-site view would predict, and it runs in the direction least convenient for anyone who assumes that blocking AI crawlers is simply good practice.
Lantad measured none of what follows. This is a reading of one paper, opened and quoted on 3 August 2026, and this site has taken no scan of any domain in its dataset. The part this site can add arrives after the figures, and it is the distinction the paper is careful to draw and most summaries of AI crawler blocking collapse: the rule a file declares and the response a server actually returns are two different measurements, and they disagree in different ways for different kinds of publisher.
| Measure | Reputable news | Misinformation | Ratio |
|---|---|---|---|
| Disallow at least one AI crawler | 60.0% | 9.1% | 6.6x |
| Average AI user agents named | 15.5 | 0.77 | 20x |
| Served a valid robots.txt | 96.4% | 73.8% | 1.3x |
| Disallow rate, September 2023 | 23% | 4.6% | 5.0x |
| Disallow rate, May 2025 | 60% | 9.2% | 6.5x |
What the study measured, and on how many sites
The paper is Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web, filed under Computers and Society, first posted on 11 October 2025 and revised twice to a third version on 16 October 2025. Its authors are Nicolas Steinacker-Olsztyn, Devashish Gosain and Ha Dao. The dataset is built from Media Bias/Fact Check classifications, which is worth stating plainly because everything downstream inherits it: reputable means a source that outlet rates Pro-Science, Least Biased, Left-Center or Right-Center with a high or very high factual reporting rating, and misinformation means one it rates Questionable Source or Conspiracy-Pseudoscience with low or very low factual accuracy. That produced 4,079 websites in total, 3,369 reputable and 710 misinformation.
Not every site in a list of that age is still answering. After filtering to homepages that were still operational, 3,299 reputable sites remained, of which 3,179 served a valid robots.txt file, an adoption rate of 96.4 percent. On the misinformation side 672 were operational and 493 of those, or 73.8 percent, served a valid file. The first asymmetry is therefore in the file itself rather than in its contents: roughly a quarter of the misinformation sites publish no usable robots.txt at all, and under the Robots Exclusion Protocol a missing file is a permissive outcome, because a status in the 400 to 499 range tells a crawler it may access any resources on the server. Absence is not neutrality. It is consent by default.
The crawler list is the other half of the method and it is larger than most people assume. The authors curated 63 distinct AI user agents from Dark Visitors, Cloudflare Radar, prior research and the ai.robots.txt repository, then fetched robots.txt from seven geographically distributed vantage points, sited in Germany, Sweden, the United States, Brazil, Africa, Asia and Australia. Sixty three names is a useful number to sit with. Every one of them is a token a site owner would have to know about, spell correctly and maintain, and the count keeps growing, which is the structural problem behind the crawler tokens that never appear in your logs. A file that names fifteen agents is not a file that has covered the field. It is a file that was accurate on the day somebody last read a list.
The longitudinal half of the study uses six Internet Archive snapshots taken in September 2023, January 2024, May 2024, September 2024, January 2025 and May 2025. That is what supports the trend figures rather than a single reading, and it is why the paper can say the gap widened rather than merely that it exists. An AI crawler policy is a moving target and a one-shot measurement cannot tell a rising line from a flat one.
Flow: MBFC credibility labels to 4,079 sites: 3,369 reputable, 710 misinformation; 4,079 sites: 3,369 reputable, 710 misinformation to Filter to operational homepages; Filter to operational homepages to Fetch robots.txt from 7 vantage points; Fetch robots.txt from 7 vantage points to 3,179 and 493 valid files; 3,179 and 493 valid files to Match against 63 AI user agents; Match against 63 AI user agents to Disallow rates and trend; 3,179 and 493 valid files to 6 Internet Archive snapshots, 2023 to 2025; 6 Internet Archive snapshots, 2023 to 2025 to Disallow rates and trend.
Which crawler names each group of sites disallows
The per-agent breakdown is where the shape of the decision shows. In the May 2025 reading, the most disallowed agent among reputable news sites is GPTBot at 52.47 percent, followed by CCBot at 48.22 percent, ChatGPT-User at 45.10 percent, Google-Extended at 43.51 percent, Anthropic-AI at 43.45 percent and ClaudeBot at 42.02 percent. The paper summarises the band as roughly half disallowing GPTBot and 40 to 50 percent restricting the other major names. Among misinformation sites in the same reading, the leader is also GPTBot, at 4.74 percent, and the rest of the top of that list is Amazonbot at 3.83 percent, PetalBot at 3.28 percent, Google-Extended and CCBot at 3.10 percent each and ClaudeBot at 2.92 percent. The paper notes that disallow rates for misinformation sites stay below 5 percent across all cases, and that over 80 percent of them disallow nothing.
Two details in the reputable numbers are worth separating. The first is that the ordering is not random: GPTBot and CCBot lead because they are the two names with the clearest public association with training corpora, and CCBot in particular is singled out in the paper for the public availability of its crawled data. That is a rational target if your concern is training rather than answer-time retrieval, and it is a distinction a site owner can only act on when a vendor publishes more than one token to act on, which most do not. Six of the nine vendors this site tracks publish exactly one, a count covered in six of the nine AI vendors we track publish one crawler token. A single token collapses training and search into one switch.
The second is the spread. The paper reports that 25 percent of reputable sites disallow more than 10 AI agents, that the median for that group exceeded 25 disallowed agents by early 2025, that the upper quartile reached nearly 50, and that the most restrictive single file names 54 distinct agents. A file naming 54 agents is not a policy anybody wrote from first principles. It is a published blocklist pasted in, and it will be stale the week a new token appears. That maintenance burden is the argument for checking what your file actually resolves to per crawler rather than trusting its length, which is what the AI crawler reference exists to enumerate.
There is also a category of client that none of these numbers reach, and the paper does not claim otherwise. Named-group rules only bind clients that look for their own name, and Google alone documents fourteen agents that a global wildcard group does not decide, which is the subject of fourteen Google agents that a robots.txt wildcard does not stop. Beyond that sit the archive crawlers whose output feeds downstream corpora, including the one behind Common Crawl's July archive. Blocking CCBot today does not retract what CCBot collected in 2023.
The declared rule and the served response are two different measurements
The most useful part of the paper for anyone running a site is the second experiment, because it stops reading files and starts sending requests. The authors picked ClaudeBot and Anthropic-AI, on the stated grounds that both are among the most commonly blocked AI crawlers and that neither has publicly disclosed IP ranges, which makes user-agent string matching the likely enforcement method. They ran a control fetch first, which succeeded with a 200-level status for 2,877 reputable and 576 misinformation sites, and after excluding non-operational domains they analysed 2,863 reputable and 565 misinformation sites.
The results do not line up with the declaration figures, and that is the finding. For Anthropic-AI, 516 reputable sites, or 18.0 percent, appear to perform active blocking, against 58 misinformation sites at 10.2 percent. For ClaudeBot the reputable figure is 711 sites at 24.8 percent and the misinformation figure is 147 sites at 25.6 percent. Read that pair again: at the response layer, for that one crawler, misinformation sites block at a marginally higher rate than reputable news sites. The six-fold gap in what the two groups declare in robots.txt does not survive into what their servers actually do to a request carrying that name.
What separates the groups instead is coherence. Of the 486 reputable sites that actively block both agents, 230 also carry a DisallowAll rule for both, which the paper puts at 47.3 percent of those blocking both and 8 percent of all reputable sites, with 4 more declaring only Anthropic-AI and 2 only ClaudeBot. On the misinformation side, exactly one site that actively blocked both had a robots.txt rule disallowing both. The reputable group's stated policy and served behaviour are related to each other. The misinformation group's blocking, where it happens, appears to come from somewhere else entirely, most plausibly generic bot mitigation applied at an edge rather than an editorial decision about AI.
This is the same class of divergence reported here before from a different direction. 234 of 592 sites that ban GPTBot in robots.txt served it a 200 measured declared bans that were not enforced; this paper measures enforcement that was never declared. Both point at the same operational fact, which is that the file and the server are separate systems maintained by separate people, and neither one predicts the other. It is also why a user agent header is evidence of nothing on its own, a point made at length in a user agent is a claim, not an identity: the paper's own limitations section says the authors do not verify crawler behaviour, and its active blocking measurement necessarily sends a string that anybody could send.
One more mechanical detail decides how a file is read before any of its rules matter. The status code your server returns for the robots.txt request itself changes the outcome, and a 404 and a 503 mean opposite things to a conformant crawler, which is set out in robots.txt, 404 and 503 are opposites. A site that intends to block everything and serves a 404 on the file has published permission.
-
Anthropic-AI, reputable18.0% 516 of 2,863 sites appeared to refuse content to a request carrying this name. -
Anthropic-AI, misinformation10.2% 58 of 565 sites, well below the reputable rate. -
ClaudeBot, reputable24.8% 711 of 2,863 sites, the higher of the two reputable figures. -
ClaudeBot, misinformation25.6% 147 of 565 sites, marginally above the reputable rate despite a 9.1 percent declaration rate.
Why blocking AI crawlers is a corpus decision, not only a site decision
The implication the authors draw is about composition. If the sites with editorial standards are withdrawing from AI crawling at 60 percent and rising, and the sites without them are withdrawing at 9 percent and barely moving, then the material that remains freely crawlable is not a random sample of the web. It is a sample tilted by credibility, in the wrong direction. The paper frames this as a growing asymmetry in content accessibility that may shape the training data available to large language models, and the hedge in that sentence is load-bearing.
It is worth being exact about what the study does and does not establish, because this is the point where a summary usually overreaches. The authors state their limitations directly in the full text: the MBFC classifications rest on human judgement and carry inherent subjectivity, they do not verify crawler behaviour, and they have no visibility into proprietary training datasets. So the paper measures what sites declare and what servers return. It does not measure what any model was trained on, and nobody outside the labs can. The claim that AI answers are therefore worse is not in the paper and should not be attributed to it. What is in the paper is that the door is open on one class of site and closing on the other.
Two things complicate even the accessibility claim. The first is retention: a corpus is cumulative, and a site that started blocking in 2024 may already be represented from 2023. Work from Common Crawl's own researchers, covered here in Common Crawl's persistent core near 40 percent, fits 51 monthly archives and puts the persistently present fraction at roughly 0.4 at domain granularity. A robots.txt edit does not reach backwards through that. The second is that robots.txt is not the only control any more, and the newer ones are not captured by counting Disallow lines. Cloudflare's managed file writes a Content-signal line whose behaviour this site examined in Cloudflare's Content-signal line asks, and the Disallow rules under it block, and the standards work meant to formalise preferences is not yet deployable, as covered in the AI preferences standard you cannot deploy yet.
For anyone doing generative engine optimization on a site that has nothing to hide, the practical reading is narrower and more useful than the ethical one. A large share of high-quality publishers have opted out of the crawlers that feed AI answers. If you have not, and your pages are actually readable, the competitive field for being the source an answer engine reaches is thinner than the raw size of the web suggests. That is not an argument for anybody to change their policy. It is a reason to know precisely which crawlers your own site currently admits, rather than assuming.
Established by the study
- 60.0% of reputable sites disallow at least one AI crawler
- 9.1% of misinformation sites do the same
- 15.5 agents named on average against 0.77
- Reputable rate rose from 23% to nearly 60%, 2023 to 2025
- Reputable served behaviour tracks declared rules more closely
Not established, and stated as such
- What any model was actually trained on
- Whether crawlers honoured the directives
- Any causal effect on answer quality
- That MBFC labels are objectively correct
- Anything about non-English sources
What to check on your own site after reading this
None of the figures above tell you anything about your domain, and the useful response to a population study is to measure the one case you control. Three checks follow from the paper's own structure, and each one is cheap.
Check what your file resolves to per crawler, not what it looks like. A robots.txt is a set of groups evaluated by name, and the outcome for GPTBot can differ from the outcome for ClaudeBot in the same file for reasons that are not visible by reading it top to bottom. Resolving a real file against a specific token is what the robots.txt tester does, and it is the difference between having a policy and having a file. Then check the second layer, because this paper is the strongest available evidence that the two disagree: send a request as a crawler and see what comes back. What GPTBot sees fetches a page the way a crawler would and shows the text that survives, and the gap between that text and what a browser renders is the failure mode prose parity is named for. A site can admit every crawler in the file and still return nothing usable, which is the case set out in two layers decide if AI can read your site.
Third, decide the training and retrieval question separately if your vendors let you, and accept that most will not. The paper's per-agent table shows publishers already treating these as different problems, blocking CCBot and GPTBot at higher rates than the retrieval-oriented names. That is only expressible where a vendor has published more than one token, and the single-token vendors force one answer to both questions.
What this site can contribute to the subject is a measurement of one domain rather than a population, and the boundaries of that are written down in the methodology. It resolves a real robots.txt against a registry of crawler tokens, fetches the page the way a crawler would, and reports what was readable and what was not, which is a component of an AI visibility score rather than a verdict on anyone's editorial policy. Aggregate findings across scans, when the sample is large enough to publish honestly, go on the research page and not into a post like this one, because this post is a reading of somebody else's work and should not be dressed as anything more.
- Resolve robots.txt per crawler token The outcome for GPTBot and ClaudeBot can differ inside one file. Reading the file top to bottom does not tell you which group binds.
- Compare the declared rule to the served response The paper found both directions of divergence. A ban that is not enforced and enforcement that was never declared are both common.
- Check the text that survives a crawler fetch Admitting a crawler in robots.txt is worthless if the page returns an empty shell without JavaScript.
- Count how many of the 63 known agents your file names The paper's most restrictive file named 54. An average reputable file named 15.5, which is not coverage of the field.
Common questions
Do reputable news sites really block AI crawlers more than misinformation sites?
In the dataset of arXiv 2510.10315, yes. The paper reports 60.0 percent of reputable news sites carrying a DisallowAll directive for at least one AI crawler in robots.txt, against 9.1 percent of misinformation sites, with average counts of named AI user agents of 15.5 and 0.77 respectively. Those figures come from 3,179 reputable and 493 misinformation robots.txt files that were valid when fetched, using Media Bias/Fact Check credibility labels to assign the two groups.
Does this mean AI models are trained mostly on misinformation?
No, and the paper does not claim it. Its stated limitations include having no visibility into proprietary training datasets, so it measures what sites declare and what servers return rather than what any model consumed. It also does not verify whether crawlers honoured the directives at all. The finding is about the asymmetry in what is left accessible, not about the contents of any training corpus.
Which AI crawler is blocked most often?
In the paper's May 2025 reading, GPTBot leads both groups: 52.47 percent among reputable news sites and 4.74 percent among misinformation sites. For reputable sites the next names are CCBot at 48.22 percent, ChatGPT-User at 45.10 percent, Google-Extended at 43.51 percent, Anthropic-AI at 43.45 percent and ClaudeBot at 42.02 percent.
Does a robots.txt Disallow actually stop an AI crawler?
It is a request, not an access control. RFC 9309 states that the protocol is not a form of access authorization, and this paper's second experiment shows the declared rule and the served response diverging in both directions: 25.6 percent of misinformation sites appeared to actively block ClaudeBot while only 9.1 percent of that group declared any AI crawler disallowed at all. Checking what a server returns to a crawler request is a separate measurement from reading the file.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.