BlogFindings
Crawl-delay in robots.txt: 152 of 1,056 sites set one, and 11 gave it to an AI crawler
Lantad requested robots.txt from all 1,419 hostnames in this repository's committed corpus on 3 October 2026 and parsed the 1,056 that returned a readable file. 152 carried at least one Crawl-delay directive and 324 directives appeared in total, but only 11 sites put one in a group naming an AI crawler token, and 237 of the 324 directives were aimed at SEO and backlink crawlers instead.
So this run counted what sites actually deploy. Lantad requested robots.txt from every hostname in this repository's committed corpus on 3 October 2026 using the scanner user agent published on its bot page, parsed each file by the grouping rules the standard sets out, and recorded every Crawl-delay found: its value, and which user-agent group it sat in. Three things came out of it. The directive is common enough to matter, it is almost never pointed at the crawlers people currently worry about, and on a measurable set of sites a parsing rule disconnects the delay from the crawler it was written for. The request discipline is the one described on the methodology page and used for every corpus measurement on this blog.
In short
- Lantad requested robots.txt from 1,419 hostnames on 3 October 2026 and parsed the 1,056 that returned a readable file. 152 of those files carried at least one Crawl-delay in robots.txt, and 324 separate directives appeared across them.
- Only 11 of the 1,056 sites put a Crawl-delay in a group naming one of the 15 AI crawler tokens this scanner tracks, although 212 of the files named at least one of those tokens somewhere in the file.
- 237 of the 324 directives sat in a group naming some other crawler, led by AhrefsBot with 39, DotBot with 24, MJ12bot with 14 and bingbot with 13, so the field is deployed mostly against SEO and backlink crawlers rather than AI ones.
- RFC 9309 has been on the Standards Track since September 2022 and never mentions crawl-delay, and Google's robots.txt documentation, carrying a last updated date of 2026-08-31, lists it as a field Google does not support. Anthropic's crawler article says it supports the non-standard Crawl-delay extension and gives ClaudeBot as its worked example.
- On 13 of the 71 sites that set a Crawl-delay in their wildcard group, at least one AI crawler token also had its own group carrying no delay, so under the group selection rule in RFC 9309 that wildcard delay never reaches that crawler.
What does Crawl-delay in robots.txt actually do?
Nothing, unless the crawler reading it has decided otherwise and said so. That is the whole answer, and it is worth stating plainly because the field looks exactly like the fields around it and behaves nothing like them.
RFC 9309 was published in September 2022 as a Standards Track document. Its grammar defines user-agent, allow and disallow, and section 2.2.4 allows that crawlers may interpret other records that are not part of the protocol, naming Sitemaps as its example. Crawl-delay is not in the grammar and is not mentioned anywhere in the document. Google's robots.txt documentation, which carried a last updated date of 2026-08-31 when it was read for this post, is more direct still: it lists user-agent, allow, disallow and sitemap as the fields Google supports, and says in the same sentence that other fields such as crawl-delay are not supported.
Among the AI crawler operators the position splits, and the split does not follow the lines a reader might expect. Anthropic's crawler article states that to limit crawling activity it supports the non-standard Crawl-delay extension to robots.txt, and the worked example it prints is a user-agent line naming ClaudeBot followed by a Crawl-delay of 1. Elsewhere on the same page it says its bots aim for minimal disruption by respecting Crawl-delay where appropriate, which is a hedge rather than a guarantee and should be read as one. The page carries no precise revision date: read on 3 October 2026 it showed only that it had been updated over six months earlier.
OpenAI's crawler documentation does not mention crawl-delay at all. Neither does Perplexity's bots guide. Both pages are detailed about access, naming their tokens and explaining which of them a robots.txt rule governs, and both are silent about rate. That silence is not the same as a refusal, and this post does not claim either operator ignores the field. It claims only what the pages say, which is nothing. Common Crawl's FAQ sits at the other end: it states that it obeys the Crawl-delay parameter for robots.txt and that increasing the number indicates to CCBot to slow its rate of crawling. Three positions, then, across five publishers of crawler documentation, and a site owner writing one directive has no way to tell from the file which of the three they are getting. Thin documentation is the normal condition in this field rather than an exception, which the earlier count of six of the nine AI vendors publishing one crawler token already showed from the access side. Access signals at least fail loudly, as the earlier finding that a 404 on robots.txt allows every crawler while a 503 blocks them all showed. A rate signal fails silently.
| Publisher | Documents Crawl-delay | What the page says |
|---|---|---|
| RFC 9309 | No | Standards Track, September 2022. Grammar defines user-agent, allow and disallow. Crawl-delay is not mentioned. |
| Not supported | Lists four supported fields and says other fields such as crawl-delay are not supported. Page updated 2026-08-31. | |
| Anthropic | Yes | Supports the non-standard Crawl-delay extension. Example names ClaudeBot with a delay of 1. |
| Common Crawl | Yes | Obeys the Crawl-delay parameter; a higher number slows CCBot. |
| OpenAI | Silent | Crawler page names its tokens and addresses access only. No mention of crawl rate. |
| Perplexity | Silent | Bots guide addresses access and user-initiated fetches only. No mention of crawl rate. |
How many sites set a Crawl-delay, and how many aim it at an AI crawler
Each of the 1,419 hostnames received one request for /robots.txt over HTTPS, from one network location, with redirects followed and a twenty second timeout. 1,106 answered HTTP 200. The rest did not, and the failure modes are worth listing because they shape what the denominator can support: 159 answered HTTP 403 to this scanner, 55 answered HTTP 503, 49 answered HTTP 404, 20 failed DNS resolution, 9 timed out, 8 answered HTTP 202, 5 answered HTTP 429, and single hosts returned 401, 406, 418, 451, 498 and 529.
Of the 1,106 that answered 200, 50 served something that was not a robots.txt file: an HTML document, or a payload carrying none of the recognised fields. Those are excluded, which leaves 1,056 readable files. This is not a random sample of the web and it should not be read as one. A site that refuses this scanner never reaches the parser, and a site that answers 403 to a robots.txt request is exactly the kind of site most likely to have opinions about crawlers, so the surviving set leans toward hosts that admit them. The same selection effect applies to every corpus measurement here and is discussed on the crawlability study page.
152 of the 1,056 files carried at least one Crawl-delay, and because a single file can carry one per group, 324 directives appeared in total. One file, france24.com, carried 22 of them. So the directive is not a historical curiosity: about one in seven of the sites that let this scanner read their robots.txt is still writing a field that the dominant search engine says it ignores and that the governing standard never defined.
Then the figure this post was written for. 212 of the 1,056 files named at least one of the 15 AI crawler tokens this scanner tracks, which is the registry behind the AI crawlers reference. Of those 212 sites, 11 gave a Crawl-delay to a group naming one of those tokens. Eleven. The sites doing it are canonical.com, coralvilleanimalhospital.com, thrivemarket.com, github.com, rivm.nl, dtu.dk, decathlon.fr, bergfreunde.de, derstandard.at, gmarket.co.kr and mountsinai.org, Because one user-agent group can name many tokens, those 15 directives carry 40 token mentions between them: GPTBot appears in six of the 15, ClaudeBot in six, OAI-SearchBot in four, ChatGPT-User, anthropic-ai, Google-Extended, PerplexityBot and Meta-ExternalAgent in three each, and the rest in one or two. One of the 15, on derstandard.at, sat in a group that also carried Disallow and a path of /, which makes the rate instruction moot: a crawler told not to fetch anything has no rate to limit. The gap between 212 sites that know the token names and 11 that set a rate for them is the finding, and it is not obviously a mistake. A site may well decide that access is the only lever worth pulling.
The directive is aimed at SEO crawlers, not AI ones
237 of the 324 directives sat in a group that named neither the wildcard nor any AI crawler token. Reading the agent names in those groups explains what the field is actually being used for, and it is not AI.
AhrefsBot led with 39 directives, followed by DotBot with 24, MJ12bot with 14, bingbot with 13, Slurp with 12, AhrefsSiteAudit with 10, Pinterest with 9, msnbot with 7, PinterestBot with 5, SemrushBot with 5, Googlebot with 4, rogerbot with 4, proximic with 4 and SwiftBot with 3. The top of that list is backlink and site-audit crawlers, the middle is general search infrastructure including two Microsoft tokens, and almost none of it is an AI crawler. The four Googlebot directives are the clearest case of effort spent on nothing, because Google's own documentation says it does not support the field.
This is what a directive looks like when it keeps being copied forward. The crawlers a Crawl-delay was written against are the ones at the top of that list, and the files carrying the field today are mostly still carrying it for them rather than for anything newer. The pattern matches what the earlier measurement of 2,209 user-agent tokens across 1,004 robots.txt files found about token lists generally, and it is the same inheritance problem that makes a renamed crawler token leave a robots.txt group matching nothing.
There is a reading of this that is not neglect, and it deserves stating because it is probably right for many of these sites. The crawlers a Crawl-delay was written against were the ones that hurt: high-frequency, low-value, commercially motivated crawls of every URL on a domain. Several AI crawlers are not that. Common Crawl archived 2.14 billion pages in a single July crawl without running any JavaScript, which is enormous in aggregate and modest per host. A site that has never seen an AI crawler cause a load problem has no reason to invent a rate limit for one, and writing a directive for every token in a registry would be the same copying behaviour that produced the Slurp lines. The honest version of this finding is not that 201 sites forgot. It is that rate control has not yet become something site owners think about per AI crawler, while access control plainly has, and anyone reasoning about crawl budget advice landing on AI crawlers is working on the access side of the same file.
| Token in the group | Directives | Operator the token itself names |
|---|---|---|
| AhrefsBot | 39 | Ahrefs |
| DotBot | 24 | Not identifiable from the token |
| MJ12bot | 14 | Not identifiable from the token |
| bingbot | 13 | Microsoft |
| Slurp | 12 | Not identifiable from the token |
| AhrefsSiteAudit | 10 | Ahrefs |
| 9 | ||
| msnbot | 7 | Microsoft |
| SemrushBot | 5 | Semrush |
| Googlebot | 4 | Google, which documents no support |
A wildcard Crawl-delay does not reach a crawler that has its own group
71 sites put a Crawl-delay in their wildcard group, which is the sensible way to set a rate for everything at once and is where 72 of the 324 directives sat. For most crawlers on most of those sites, it does what it looks like it does. On 13 of them it does not, and the reason is a single sentence of the standard.
RFC 9309 says that crawlers must use case-insensitive matching to find the group matching their product token and then obey the rules of that group, and that if more than one group matches, the matching groups must be combined into one. Then the sentence that matters here: if no matching group exists, crawlers must obey the group with a user-agent line with the asterisk value, if present. The wildcard group is the fallback, not a baseline. A crawler that finds its own name in the file reads its own group and stops looking, and anything written only in the wildcard group is invisible to it.
That is a well-known trap for Disallow lines, and it is the same trap for Crawl-delay with one difference: a lost Disallow is usually noticed eventually, because content shows up somewhere it should not. A lost rate limit is noticed only as load. Of the 71 sites with a wildcard Crawl-delay, 15 also named a tracked AI crawler token somewhere in the file, and on 13 of those at least one named token had its own group that carried no Crawl-delay. Counting across tokens, CCBot was shadowed this way on 8 of those sites, ClaudeBot on 7, Meta-ExternalAgent on 7, Bytespider on 7, GPTBot on 5, anthropic-ai on 5, Applebot-Extended on 5 and Amazonbot on 5.
The sites are worth naming because the pattern is consistent and none of it looks careless. congress.gov sets a wildcard delay of 2 and gives its own group to 14 tracked tokens, none of them carrying a delay. scientificamerican.com sets 5 and gives its own group to 10. chronicle.com sets 10 and gives its own group to 9. healthline.com sets 5 and gives its own group to 8. villagedentaldtc.com sets 1 and gives its own group to 9. These are files written by someone who knew the token names, went to the trouble of listing them, and in doing so moved those crawlers out of the group holding the rate instruction. Whether that is a problem depends entirely on which crawler it is, because only the operators that document honouring the field would have slowed down for it anyway, and the earlier teardown of two layers deciding whether AI can read your site applies here too: the file is the weaker layer, and a rate instruction inside it is the weakest thing in the file. The robots.txt tester resolves a given token against a given file by these rules, which is the only reliable way to tell which group a named crawler actually lands in.
Flow: Crawler reads robots.txt to Own product token in a group?; Own product token in a group? (yes) to Obey that group only; Own product token in a group? (no) to Fall back to wildcard group; Obey that group only to Wildcard Crawl-delay not applied; Fall back to wildcard group to Wildcard Crawl-delay applied.
What the values say, from 0 seconds to 1,000
All 324 directives carried a numeric value. Not one was malformed, which is a mildly surprising result for an unstandardised field and suggests the directive is being copied from working examples rather than improvised.
The median was 10 seconds and 10 was also the single most common value, on 152 directives, very nearly half the total. After that the distribution falls away quickly: 1 second on 41, 5 seconds on 40, 3 on 24, 2 on 18, 60 on 12, 30 on 9 and 20 on 8. The long tail holds a handful of fractional values, with 0.5 on two directives, 0.2 on two and 0.1 on one, and two directives set 0, both on mcgill.ca and both for enterprise search crawlers rather than public ones. A delay of 0 is a statement that no delay applies, which is also what the absence of the field means, so those two are decoration.
The top of the range is where the interesting case sits. dtu.dk, the Technical University of Denmark, sets a Crawl-delay of 10 in its wildcard group, then repeats its entire rule set in a group for ClaudeBot and again in a group for GPTBot, and in both of those groups raises the delay to 1,000. A third group does the same for GoogleOther. A thousand seconds between requests is a little under 87 requests a day, which for a university site of any size means a complete crawl would take months. This is the only site in the corpus that aimed a deliberate, order-of-magnitude throttle at named AI crawlers, and it did it the way Anthropic's documentation asks: the delay sits inside the token's own group, so the group selection rule works for it rather than against it.
Whether both halves of that throttle work is the asymmetry this whole post keeps running into. Anthropic documents supporting the field, so ClaudeBot plausibly honours a delay of 1,000. OpenAI's documentation does not mention the field, so what GPTBot does with the same line is not knowable from the file or from the docs, and this post makes no claim about it. The site wrote one instruction and may be getting two different behaviours. That is the structural problem with rate control in robots.txt and the reason access directives remain the only part of the file worth relying on, which is also why 234 of 592 sites that ban GPTBot in robots.txt still served it a 200 is a finding about enforcement rather than about syntax.
What to check on your own robots.txt
Four checks follow from the measurement, and none of them needs a tool more complicated than reading the file carefully in the right order.
First, work out which group each crawler you care about actually lands in. This is the check that catches the 13-site pattern above, and it is the only one of the four that is a correctness question rather than a judgement call. If a token has its own group, every rule you want applied to it has to be inside that group, repeated, including the rate. Inheritance does not happen. The same resolution logic decides which Disallow lines apply, so getting it wrong costs more than a lost delay, and it is worth running against whichever tokens matter to you rather than reasoning about it from the file by eye.
Second, decide whether you want a rate limit at all, separately from whether you want access. These are different questions and the corpus suggests most sites have answered only the second. 212 files named an AI crawler token and 11 set a rate for one. If your origin has never struggled under crawler load, the honest answer is that you do not need the field, and adding it to look thorough adds a line that most crawlers will ignore.
Third, if you do want one, put it where the operator's own documentation says to put it and expect it to work only for the operators that document it. That currently means Anthropic's crawlers and CCBot among the tokens measured here. For Google it is documented not to work. For OpenAI and Perplexity there is no documentation either way, which means you are guessing, and a rate limit you are guessing about is not a control.
Fourth, treat the whole file as advisory and verify the outcome rather than the instruction. Rate is the softest signal in a file whose hardest signals are routinely not enforced, and an edit to robots.txt does not take effect when you save it because crawlers cache the file. What you can verify is what a crawler is served, which is what what GPTBot sees reports, and what your logs say arrived. Lantad did not measure crawl rates for this post and cannot: rate is visible only in a site's own server logs over time, and nothing in a single robots.txt request shows whether any crawler honoured any delay. What is measured here is what 1,056 sites wrote down, which is a claim about deployment and not about compliance.
- Resolve the group for each token you care about A crawler finding its own group ignores the wildcard group entirely, rate included. This is the error found on 13 of 71 sites.
- Separate the rate question from the access question 212 corpus files named an AI token, 11 set a rate for one. Most sites need only the access answer.
- Check the operator documents the field Anthropic and Common Crawl document honouring it. Google documents not supporting it. OpenAI and Perplexity say nothing.
- Verify what is served, not what is written Crawlers cache robots.txt, and rate compliance appears only in server logs over time. Not measurable from the file.
Lantad
Published .
A robots.txt file has three fields that every crawler operator agrees about and one that nobody ever standardised. User-agent, allow and disallow decide access, and RFC 9309 puts them on the Standards Track with a grammar. Crawl-delay decides rate, and it appears in no standard at all. That gap matters more than it used to, because the question a site owner now asks about an AI crawler is often not whether it may read a page but how hard it may hit the origin while doing it. Access is a yes or a no. Rate is a negotiation conducted in a field with no referee.
Common questions
Is Crawl-delay part of the robots.txt standard?
No. RFC 9309, published on the Standards Track in September 2022, defines user-agent, allow and disallow, and does not mention crawl-delay anywhere. Its section 2.2.4 permits crawlers to interpret records outside the protocol and names Sitemaps as the example, which is the clause an unstandardised field like Crawl-delay relies on.
Do AI crawlers honour Crawl-delay in robots.txt?
It depends on the operator and the only honest source is each operator's own documentation. Anthropic states it supports the non-standard Crawl-delay extension, and Common Crawl states it obeys the parameter for CCBot. Google's documentation, updated 2026-08-31, says the field is not supported. OpenAI's and Perplexity's crawler pages do not mention it, so neither honouring nor ignoring can be claimed for them.
Why does my wildcard Crawl-delay not apply to GPTBot?
Because RFC 9309 makes the wildcard group a fallback rather than a baseline: a crawler obeys the group matching its own product token, and only falls back to the asterisk group if no group matches it. If your file gives GPTBot its own group, every rule you want applied to GPTBot has to be repeated inside that group. Lantad found this pattern on 13 of the 71 corpus sites that set a wildcard Crawl-delay.
What Crawl-delay value do sites actually use?
Across the 324 directives Lantad found on 3 October 2026 the median was 10 seconds, and 10 was the most common value on 152 of them. 1 second appeared on 41 directives and 5 seconds on 40. The maximum was 1,000 seconds, set by dtu.dk for ClaudeBot and GPTBot specifically, and the minimum was 0, which means the same as writing nothing.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.