BlogFindings

AI crawler statistics: 77 of 706 robots.txt files blocked a citation crawler, and 42 of those were news sites

Lantad requested robots.txt and the home page from all 1,027 hostnames in this repository's committed industry corpus seed file on 29 September 2026 as LantadBot. 706 returned a parseable robots.txt. 77 of those disallowed GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot at the root, and 42 of the 77 came from the news stratum alone.

17 min read Lantad

Lantad requested /robots.txt and the home page from all 1,027 hostnames in the industry corpus seed file committed to this repository, on 29 September 2026, as LantadBot, following redirects and executing no JavaScript. The frame is stratified by industry: government, education, healthcare, news, SaaS, ecommerce, travel and finance, roughly 130 hostnames each. The result is that the widely repeated claim that the web is closing to the AI crawler is, on this corpus, a claim about one sector. Seven of the eight have barely written a rule.

In short

  • Lantad requested /robots.txt and the home page from all 1,027 hostnames in this repository's committed industry corpus seed file on 29 September 2026 as LantadBot/1.0: 706 returned a parseable robots.txt, and 156 of those named at least one of 31 crawler tokens their operators tie to AI.
  • These AI crawler statistics describe newsrooms rather than the web: 77 of the 706 files disallowed GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot at the root on 29 September 2026, and 42 of those 77 were news sites.
  • Outside news the rate collapses. On 29 September 2026, 2 of 92 finance files, 2 of 76 travel files and 3 of 119 SaaS files blocked any of those four crawlers at the root, against 42 of 64 news files.
  • 49 of the 72 files that disallowed GPTBot at the root still allowed OAI-SearchBot on 29 September 2026, and 23 of those 49 were news sites, so the common decision is to refuse the training crawler and keep the search crawler.
  • 325 of the 1,027 home pages returned no readable HTML to LantadBot on 29 September 2026, and only 39 of those 325 answered a browser user agent from the same machine, so most of that refusal was not a decision about the user agent string.
StepHostnamesNote
Hostnames requested1,027Eight industry strata, 123 to 130 each
Returned a parseable robots.txt706HTTP 200 and a body that was not HTML
Answered robots.txt with a web page171An HTML document where a text file was asked for
Refused robots.txt, no HTML body11959 sent HTTP 403, 56 sent 503 or similar, 4 sent 404
Did not complete a robots.txt request31DNS, TLS or a twenty second timeout
Named at least one AI crawler token156 of 70622.1 percent of the files that parsed
Disallowed a citation crawler at the root77 of 706GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot
Of those 77, drawn from the news stratum42One of eight strata supplies 54.5 percent of them
Home page returned readable HTML697325 returned an error, a challenge or nothing
What Lantad found when it requested /robots.txt and the home page from all 1,027 hostnames in this repository's committed industry corpus seed file on 29 September 2026 as LantadBot/1.0, redirects followed and no JavaScript executed. Percentages are of the 1,027 hostnames requested.

AI crawler statistics: what did 1,027 sites declare on one day?

706 of the 1,027 hostnames answered a request for /robots.txt with HTTP 200 and a body that was not an HTML document. The 321 that did not are not a rounding error and they are not all failures of the same kind. Counted so that each hostname appears once: 171 answered the request for a text file with an HTML document, 59 refused with HTTP 403 and no HTML body, 56 answered HTTP 503 or a near neighbour, 4 returned HTTP 404 as text, and 31 did not complete a connection inside twenty seconds. Counted by status instead, 160 of the 321 carried an HTTP 403 and 101 of those 160 also returned HTML, which is why the two counts have to be reported separately rather than added. A file that comes back as HTML is the defect this site has counted before, where 232 of 1,059 robots.txt files carried a defect and 33 sites answered with a web page, and our robots.txt tester exists to show a site owner which of those they are serving.

Each of the 706 files was parsed the way RFC 9309 specifies, because a robots.txt figure is only as good as the matching rule behind it. The RFC says a crawler must use case-insensitive matching to find the group matching its product token, must obey the group with a user-agent line of "*" if no group matches, and that where several rules match a path the most specific match, the one with the most octets, must be used. So a site that writes a Disallow under User-agent: * and never names GPTBot is counted here as blocking GPTBot, and a site that names GPTBot in an empty group is counted as allowing it.

156 of the 706 files named at least one of 31 tokens that their operators tie to AI training, AI search or AI agents. That list includes Applebot and Bingbot, which are search crawlers whose operators also run answer features, so the 156 is a count of files that have considered the subject at all rather than a count of blocks. The median file in that group named 7 tokens and the longest named 28, which is the length at which a list stops being a decision and starts being a paste. Naming a token is also not the same as matching one, and this site has counted how far that goes: on 17 September 2026, 1,004 robots.txt files named 2,209 distinct tokens between them, 813 of which contain a character RFC 9309 does not let a crawler put in its own name.

The per-sector shape of that naming is its own finding, and it is not the same shape as the blocking. Government and travel both have a median of 1 named token among the files that name any, which is a single line about a single crawler. Finance has the highest median of the eight at 17 tokens, across only 6 files out of 92 that name anything at all, so in that sector the rare site that addresses AI crawlers addresses almost all of them at once and then, in 4 of those 6 cases, blocks none of the four measured here.

SectorHostsParsedNamed a tokenBlocked a citation crawler
News128644542
Ecommerce13083287
Education130103179
Healthcare12388148
SaaS130119233
Government12981124
Travel13076112
Finance1279262
Eight industry strata measured by Lantad on 29 September 2026. Parsed is the count of hostnames returning HTTP 200 and a non-HTML body for /robots.txt. Named and Blocked are counted out of that parsed count, never out of the stratum size, because a file that never arrived states no policy.

Which sectors block an AI crawler, and which never wrote a rule?

77 of the 706 parsed files disallowed at least one of GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot at the root. 42 of those 77 came from the news stratum, which contributed 64 of the 706 files. Put the other way round: 65.6 percent of the news files that parsed block a citation crawler, against 2.2 percent of finance, 2.6 percent of travel and 2.5 percent of SaaS. That is not a difference of degree between sectors that are all moving in the same direction. It is one sector that has made a decision and seven that have not been asked to.

The named sites make the split legible. In news, bloomberg.com, cnn.com, cbsnews.com, aljazeera.com, abc.net.au, bostonglobe.com, forbes.com and irishtimes.com all disallow GPTBot at the root. In finance the entire blocking set across 92 parsed files is ally.com and bancolombia.com.br, and in travel across 76 files it is airbnb.com and tripadvisor.com. This is the same direction of travel an arXiv study found when it compared reputable news against misinformation sites, which this site read and reported in full: 60.0 percent of reputable news sites disallowed at least one AI crawler against 9.1 percent of misinformation sites. Two different frames, two different methods, one conclusion about who writes these rules.

Naming a crawler is not the same as blocking it, and the gap is the second finding here. 81 of the 156 files that named a token blocked none of the four, among them cloudflare.com, cockroachlabs.com, bestbuy.com, chewy.com, booking.com, expedia.com, gov.uk and census.gov. Those files typically name a long list and disallow something else entirely, most often a training-only or scraper token such as CCBot or Bytespider. A site owner reading a pooled blocking rate as evidence that generative engine optimization is under threat in their own market should first check which market the rate came from, because this site has also measured GPTBot named in 82 of 718 robots.txt files as the most blocked AI crawler without being able to say, until now, where those 82 were concentrated.

  • News 65.6% 42 of 64 files
  • Education 8.7% 9 of 103 files
  • Healthcare 9.1% 8 of 88 files
  • Ecommerce 8.4% 7 of 83 files
  • Government 4.9% 4 of 81 files
  • SaaS 2.5% 3 of 119 files
  • Travel 2.6% 2 of 76 files
  • Finance 2.2% 2 of 92 files
Percentage of each sector's parsed robots.txt files that disallowed GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot at the root, measured by Lantad on 29 September 2026. The denominator is the files that parsed in that sector, shown in the note.

Why 49 of 72 sites that block GPTBot still allow OAI-SearchBot

72 of the 706 files disallowed GPTBot at the root. 49 of those 72 allowed OAI-SearchBot at the same time, and 23 of the 49 were news sites. The same asymmetry runs through the Anthropic pair in the other direction: 65 files disallowed ClaudeBot and 24 of them allowed PerplexityBot. These are not inconsistencies. They are a policy that the vendor documentation makes available and that most published advice still flattens into one question about whether to allow AI.

OpenAI's crawler documentation, opened for this post on 29 September 2026, separates four agents and carries no date of its own. It says OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features", and that GPTBot "is used to crawl content that may be used in training our generative AI foundation models". Those are different products behind different tokens, and OpenAI publishes them on one page. Anthropic's crawler article, last updated 7 April 2026, describes ClaudeBot as collecting web content that could contribute to training, and documents Claude-User and Claude-SearchBot separately; Anthropic's own page is the source for that. Perplexity's bot guide, which carries no date, states that PerplexityBot "is designed to surface and link websites in search results on Perplexity" and "is not used to crawl content for AI foundation models".

So the 49 files are doing exactly what the three vendors describe: refusing the training crawler and keeping the crawler that can put a link in an answer. Whether that intention survives contact with the token list is a separate question, and this site has measured the ways it fails. 17 of the 64 news files disallow OAI-SearchBot as well, which ends the citation path they presumably wanted to keep, and the same shape shows up where 38 robots.txt files refuse GPTBot at the root and let ChatGPT-User in. Naming the search token at all is still unusual: this site counted OAI-SearchBot in 16 of 140 robots.txt files while 25 named GPTBot. If the goal is to appear in an answer rather than in a model, the rules that achieve it are the ones set out on how to get cited by ChatGPT.

  • Disallowed GPTBot 72 of 706 OpenAI's training crawler, and the most blocked token of the four.
  • Disallowed ClaudeBot 65 of 706 Anthropic's training crawler. 24 of the 65 allow PerplexityBot.
  • Disallowed PerplexityBot 43 of 706 Perplexity documents it as a search crawler, not a training one.
  • Disallowed OAI-SearchBot 24 of 706 The token that decides ChatGPT search visibility, blocked least often.
  • Blocked training, allowed search 49 of 72 Files disallowing GPTBot while allowing OAI-SearchBot. 23 are news sites.
  • Named a token, blocked none of four 81 of 156 A list of AI crawlers in the file is not evidence of a block.
How the 706 parsed robots.txt files ruled on four citation crawlers at the path /, measured by Lantad on 29 September 2026 under RFC 9309 matching, so a site-wide Disallow under User-agent: * counts as a block for every token that has no group of its own.

325 home pages did not answer, and robots.txt said nothing about it

The robots.txt figures describe policy. They say nothing about whether a crawler gets the page, and on this corpus the two answers disagree badly. 325 of the 1,027 home pages returned no readable HTML to LantadBot on 29 September 2026. Counted once each: 218 answered HTTP 403, 56 answered HTTP 503, 7 answered HTTP 429, 12 answered some other 4xx or 5xx, 31 did not complete a connection, and 1 returned HTTP 200 with an interstitial challenge body instead of a page. Cutting the same 325 by body rather than by status, 105 carried an interstitial challenge, so for the great majority the challenge page and the error status are one refusal counted two ways. Almost none of it is written down in a robots.txt file anywhere.

The obvious reading is that these sites refuse crawlers by user agent, and the control says otherwise. Every one of the 325 was asked again from the same machine, seconds later, with a current desktop Chrome user agent string. Only 39 of the 325 came back with HTTP 200 and at least a hundred words. The other 277 refused the browser string too. So for five out of six of these hostnames the refusal is not a judgement about the user agent at all: it is an address, an autonomous system, a TLS fingerprint or a rate, and changing the name the crawler gives would not move it. That is inconvenient for a product that reports what a named crawler sees, and it is the honest limit of any single-machine measurement including this one.

The 39 that did turn on the user agent are the ones a site owner can actually fix, and they are spread evenly rather than concentrated: 9 travel, 8 ecommerce, 7 government, 6 finance, 3 healthcare, 2 each in news, education and SaaS. This site has measured the same shape before from the other end, where 79 of 115 sites refused a crawler their own robots.txt allowed, and where 103 of 1,089 sites served an unknown bot and refused GPTBot. One caution belongs on the news figures because of this: 58 of the 128 news hostnames refused the robots.txt request itself, so the 42 blocks are counted from the 64 newsrooms that let the file through, and the ones that did not are unlikely to be the permissive half. The news rate here is more likely a floor than a ceiling. If you want to see which of these applies to your own domain, what GPTBot sees runs the fetch rather than reasoning about it.

Refused the crawler and the browser: 277

  • The same response for both user agent strings
  • 218 of the 325 answered HTTP 403
  • 56 answered HTTP 503 and 7 answered HTTP 429
  • 105 of the 325 carried a challenge body at any status
  • Address, network or fingerprint, not the name
  • Renaming the crawler would change nothing

Served the browser only: 39

  • HTTP 200 and at least 100 words for Chrome
  • Travel 9, ecommerce 8, government 7, finance 6
  • Healthcare 3, news 2, education 2, SaaS 2
  • This is user agent filtering and it is fixable
  • None of the 39 declares it in robots.txt
  • 12 percent of the refusals, not the majority
The 325 home pages that returned no readable HTML to LantadBot/1.0 on 29 September 2026, each then re-requested once from the same machine with a current desktop Chrome user agent string. A control that succeeds isolates user agent filtering; a control that fails points at the network layer instead.

What an AI crawler actually got from the 697 sites that answered

697 hostnames returned HTTP 200 with a body that was not a challenge. Counting the words left after script, style, noscript and comment content is removed and tags are stripped, the median home page in that set carried 1,108 words to a client that executes no JavaScript. The sector spread is wide and it does not follow the blocking spread at all. News is the longest at a median of 2,099 words, then SaaS at 1,445 and finance at 1,338, while government sits lowest at 655 and travel at 833. The sector most likely to refuse a crawler is also the sector that hands the most text to the ones it lets in.

The floor matters more than the median. 60 of the 697 home pages carried fewer than 50 words in raw HTML, which for a crawler that does not render is close to an empty document. Travel contributed 13 of those, ecommerce 10, healthcare 9, finance 9, education 8 and government 8, against 2 in SaaS and 1 in news. That is the prose parity failure this scanner grades first, and the count runs in the same direction as the 17 of 380 home pages that sent a crawler zero words on this repository's platform frame on 11 September 2026, though that run used a stricter test of no visible text at all.

Structured data was present in raw HTML on 377 of the 697, which is 54.1 percent, and the sector order changes again: 54 of 60 news pages carried a JSON-LD block against 27 of 87 government pages and 32 of 104 education pages. Public sector and university home pages are therefore the two places in this corpus where a machine gets neither much prose nor an explicit entity declaration, which is a worse combination than either alone. It is also the same direction as the 141 of 382 home pages that carried no structured data in the raw HTML measured on this repository's platform frame on 11 September 2026, and it is why a report that grades access without grading content would call a government home page healthy.

SectorPagesMedian wordsUnder 50 wordsCarried JSON-LD
News602,099154
SaaS1181,445289
Finance941,338953
Healthcare931,037943
Education104891832
Ecommerce658661041
Travel768331338
Government87655827
What LantadBot/1.0 received from the 697 home pages that answered HTTP 200 with a page rather than a challenge on 29 September 2026. Words are counted from raw HTML with script, style, noscript and comments removed and no JavaScript executed, so the figure is what a non-rendering crawler reads.

What to check on your own site, and what this does not show

The order to work in follows from which of the three layers is failing, and they fail independently. First establish that the file arrives: 321 of 1,027 hostnames could not return a parseable robots.txt to a plain text request, and a file served as HTML states no policy however carefully it was written. Second, read the rules against the tokens you actually care about rather than against the word AI, because the training and search crawlers are separate products and 49 sites in this corpus have already drawn that line deliberately. Third, fetch your own home page with a crawler-shaped client and count the words that survive, because policy and access answered differently here on hundreds of hosts.

What this run cannot show is worth stating at the same length. It is one request per file per host on one day from one machine on one network, so a site that refused us may serve a different network, and a rule that changed on 30 September is not tracked. The corpus is a stratified sample of 1,027 hostnames drawn on 3 August 2026 to study crawlability, not a random sample of the web, so every rate here describes this frame; our research says how that frame was built. The 403 and 503 responses are counted as refusals without any claim about who configured them, since a scanner cannot tell a deliberate bot rule from a CDN default. And the browser control ran only against the 325 that failed, so no claim is made about how the 697 successes would have differed.

The largest limit is the one the product itself is honest about. None of this measures whether an answer engine cites you. It measures whether the door is open, which is the input, and our methodology is explicit that access and prose are graded while placement in an answer is not. AI visibility work that starts with the ranking and works backwards keeps arriving at these same three layers, and on 29 September 2026 the layer that failed most often across 1,027 sites was not the one anybody writes rules about.

The three layers this run measured separately, and the order to check them in. Each layer can pass while the next fails, which is why 706 files stating a policy and 697 pages answering a crawler are not the same 697 sites.

Written by

Lantad

Published .

Almost every measurement this site publishes pools the whole corpus and reports one number for the web. That hides the thing a site owner most wants to know, which is whether the behaviour being described is normal in their own field or normal three fields over. So this run kept the strata apart and asked the same questions of eight of them.

Common questions

What share of websites block AI crawlers?

On this corpus, 77 of the 706 hostnames that returned a parseable robots.txt on 29 September 2026 disallowed GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot at the root, which is 10.9 percent. That pooled figure is misleading on its own because 42 of the 77 were news sites. Measured inside each stratum the rate runs from 65.6 percent of news files down to 2.2 percent of finance files.

Which industries block AI crawlers the most?

News, by a wide margin. 42 of the 64 news robots.txt files that parsed on 29 September 2026 disallowed at least one of the four citation crawlers, against 9 of 103 in education, 8 of 88 in healthcare, 7 of 83 in ecommerce, 4 of 81 in government, 3 of 119 in SaaS, 2 of 76 in travel and 2 of 92 in finance. 58 of the 128 news hostnames refused the robots.txt request itself, so the news rate is more likely a floor than a ceiling.

Why do sites block GPTBot but allow OAI-SearchBot?

Because they are different products with different tokens, which OpenAI's own crawler documentation states: OAI-SearchBot surfaces websites in ChatGPT's search features, while GPTBot crawls content that may be used in training foundation models. 49 of the 72 files that disallowed GPTBot at the root on 29 September 2026 allowed OAI-SearchBot, and 23 of those 49 were news sites, so refusing training while keeping the citation path is the common choice rather than an inconsistency.

How was this measured?

Lantad requested /robots.txt and the home page from all 1,027 hostnames in this repository's committed industry corpus seed file on 29 September 2026 as LantadBot/1.0, following redirects, executing no JavaScript, with a twenty second timeout on the file and twenty five seconds on the page. Rules were evaluated under RFC 9309 matching at the path /. Every home page that failed was re-requested once with a desktop Chrome user agent string as a control.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.