BlogFindings
Block ChatGPT: 38 robots.txt files refuse GPTBot at the root and let ChatGPT-User in
Lantad asked all 1,419 hostnames in this repository's committed corpus for robots.txt once each on 28 September 2026 and evaluated the site root for every OpenAI token with the scanner's own parser. 1,025 files parsed into at least one group. 172 named GPTBot, 103 named ChatGPT-User, 80 named OAI-SearchBot and 4 named OAI-AdsBot, and on 38 files a rule refuses GPTBot at the root while ChatGPT-User is allowed there.
OpenAI's crawler documentation, read at source on 28 September 2026, describes four agents. GPTBot collects training data. OAI-SearchBot powers search. OAI-AdsBot checks landing pages submitted as ads. The fourth is ChatGPT-User, which fetches a page because somebody in a chat window asked a question that needed it, and the page says of it that because these actions are initiated by a user, robots.txt rules may not apply. That sentence is the whole subject of this post. This run counted how many sites have written a rule for each of the four, and how often a file that refuses the training crawler leaves the user-initiated one alone.
In short
- Attempts to block ChatGPT with robots.txt have to contend with OpenAI's own documentation, read at source on 28 September 2026, which says of ChatGPT-User that because these actions are initiated by a user, robots.txt rules may not apply, and which directs opt-outs for automatic crawling to OAI-SearchBot instead.
- OpenAI documents four agents. Of the 1,025 robots.txt files that parsed into at least one group on 28 September 2026, 172 named GPTBot, 103 named ChatGPT-User, 80 named OAI-SearchBot and 4 named OAI-AdsBot, and 77 of the 172 files naming GPTBot named ChatGPT-User nowhere.
- On 38 of those 1,025 files a rule refuses GPTBot at the site root while ChatGPT-User is allowed there. The wildcard group decided for ChatGPT-User on 24 of the 38, a group naming ChatGPT-User itself decided on 12, and no group applied at all on 2. Only 2 files, cambridge.org and gmarket.co.kr, run the other way.
- Three vendors ship a fetcher for user-initiated requests and document it three different ways: Perplexity's crawler page states that since a user requested the fetch, this fetcher generally ignores robots.txt rules, while Anthropic's page dated 7 April 2026 says its bots respect do not crawl signals by honoring industry standard directives in robots.txt.
- Lantad sent no request as any OpenAI agent and observed no OpenAI fetch. Every figure here is a statement about published permission on 28 September 2026, read from the file the site serves, and not about what any vendor's software did.
| Token | What OpenAI says it does | What the page says about robots.txt | Files naming it |
|---|---|---|---|
| GPTBot | Collects content for training foundation models | Disallowing it signals content should not be used in training | 172 |
| OAI-SearchBot | Surfaces sites in ChatGPT search results | Recommends allowing it; the token to use for search opt-outs | 80 |
| ChatGPT-User | Visits a page for a user action in ChatGPT or a Custom GPT | Robots.txt rules may not apply | 103 |
| OAI-AdsBot | Validates pages submitted as ads on ChatGPT | Nothing stated | 4 |
Can you block ChatGPT with robots.txt?
Partly, and the part that works is the part nobody argues about. Three of the four OpenAI agents are ordinary crawlers in the sense that RFC 9309 describes, which is software that decides for itself what to fetch. For that kind of client the standard is a clean contract: the crawler asks for the file, selects the group whose product token matches its own name, and the longest matching rule decides. A Disallow line aimed at GPTBot is a statement that model training should not use the page, and OpenAI's documentation says exactly that.
ChatGPT-User is not that kind of client. It fetches because a person typed a question and the answer needed a page, which puts it in a category the robots exclusion standard was never written for. The standard governs automated crawling, and a fetch made once, in response to one instruction from one human, is closer to that human opening the page in a browser than to a crawler working through a site. OpenAI's page makes the same distinction in its own words, saying that ChatGPT-User is not used for crawling the web in an automatic fashion and directing anyone managing opt-outs for automatic crawling to OAI-SearchBot instead.
Whether you find that reasonable is a separate question from whether it changes what your file does, and it does change it. A site that writes one Disallow for GPTBot has opted out of training. It has not necessarily opted out of appearing in an answer, because the agent that fetches for the answer is a different token with a documented caveat attached. This is the gap that makes the popular advice misleading rather than wrong, and it is the reason our own guide to being cited in ChatGPT treats the two cases separately instead of printing one line.
It is worth being precise about what any of this evidence can settle. Published permission is not behaviour, and the two have been measured apart before: a controlled study of ten assistants found that six of them never requested robots.txt at all in any of its trials. What this scanner reads is the permission a site publishes for a named AI crawler, which is a fact about the site rather than a prediction about the vendor.
Sample Illustrative, not a measurement of any real site.
Flow: User asks ChatGPT a question to ChatGPT-User fetches the page; ChatGPT-User fetches the page to Robots.txt rules may not apply; OpenAI crawls for search or training to OAI-SearchBot or GPTBot; OAI-SearchBot or GPTBot to Robots.txt group decides.
How many robots.txt files write a rule for ChatGPT-User?
Lantad requested /robots.txt once over HTTPS from each of the 1,419 hostnames in the two committed corpus seed files on 28 September 2026, as LantadBot with redirects followed and a fifteen second timeout, then parsed every answer with the scanner's own parser. 43 hostnames produced no status at all. 261 answered a non-200. 1,115 answered HTTP 200, of which 38 returned HTML rather than robots text, leaving 1,077 plain-text files of which 1,025 parsed into at least one user-agent group. Every figure below is out of those 1,025.
208 of the 1,025, just over a fifth, name at least one of the fifteen AI crawler tokens this scanner evaluates. Inside that fifth the attention is lopsided. 172 files name GPTBot and 103 name ChatGPT-User, so the training crawler is named on two thirds more files than the agent that fetches for answers. 77 of the 172 files that name GPTBot do not mention ChatGPT-User anywhere. Going the other way is rare: 8 files name ChatGPT-User without naming GPTBot, among them cambridge.org, circleci.com, instacart.com, expedia.com and kayak.com.
The same asymmetry shows up as a count of sites that wrote AI rules and skipped the class entirely. 102 of the 1,025 files name at least one of the nine training tokens and name none of the three user-initiated fetchers. That is not carelessness so much as the shape of the advice: the blocking guides that circulated through 2024 and 2025 were written about training, and the token lists in them have outlived the question they answered. We found the same drift when we counted every token in the corpus and found 2,209 distinct names across 1,004 files, many of them products that no longer exist under that name.
The 103 files that do name ChatGPT-User do not agree about what to say to it. 41 give it Disallow and a bare slash, refusing the whole site. 37 refuse some paths and allow the rest. 24 carry no Disallow line at all inside the group, which under the standard is an explicit permission rather than an omission, and one group holds no rules whatever. So of the sites paying attention to this token, roughly a quarter are using it to invite the fetch rather than to stop it. Our earlier counts of GPTBot in 82 of 718 files and of OAI-SearchBot in 16 of 140 were drawn on smaller corpora, and the ranking between the OpenAI tokens has held across all three. The full set of tokens this scanner evaluates is published on the AI crawlers page.
Where a GPTBot block leaves ChatGPT-User allowed
Counting names is not the same as counting outcomes, because a file can refuse a token it never names. So the scanner evaluated the site root for each OpenAI token against every file, using the same group selection and longest-match rules the standard defines. At the root, 81 of the 1,025 files refuse GPTBot and 45 refuse ChatGPT-User. The interesting number is the overlap: on 38 files a rule refuses GPTBot at the root while ChatGPT-User is allowed there.
How those 38 came about matters more than the total, because two different mistakes produce the same result. On 24 of the 38 the wildcard group decided for ChatGPT-User, which means the site wrote a named rule for GPTBot, wrote nothing for ChatGPT-User, and the general group it fell through to permits the root. On 2 files no group applied to ChatGPT-User at all, which the standard treats as allowed. Those 26 are the accident. The other 12 are a decision: a group naming ChatGPT-User itself decided the verdict, and decided to allow it, on sites including figma.com, svt.se, theregister.com, ebay.com, instacart.com, airbnb.com and tripadvisor.com. A publisher refusing training while keeping user-initiated retrieval is a coherent position, and on those 12 files it is written down rather than inferred.
Only 2 files in the corpus run the other way, allowing GPTBot at the root while refusing ChatGPT-User: cambridge.org and gmarket.co.kr. That ratio of 38 to 2 is the finding. Where the two OpenAI tokens diverge in this corpus, the training crawler is almost always the one that loses, and the agent that fetches a page to answer somebody's question is almost always the one that gets through.
Widening from the root to every path the file itself mentions makes the pattern larger rather than different. Taking each file's own Disallow and Allow patterns as candidate paths, 55 files give GPTBot and ChatGPT-User different answers somewhere, and on 47 of those 55 there is at least one path GPTBot loses and ChatGPT-User keeps. This is the same structural problem we measured from the other end when 559 of the 581 paths GPTBot lost were closed by a rule that never named it: the wildcard group does nearly all the work in these files, and which token a rule names is usually incidental to what happens. The mechanics of that evaluation are set out in the methodology, and the corpus is an editorial sampling frame rather than a random draw, so these are rates for these 1,419 hostnames and nothing wider.
One more caveat belongs here, because it cuts against the numbers above. A file can only govern a crawler that reads it, and 92 of 1,056 sites in an earlier sweep refused GPTBot the robots.txt itself. A site whose infrastructure will not serve the file to a declared bot has not configured a policy that bot can follow, whatever the file says, which is a failure mode worth understanding before anyone treats robots.txt as the whole of a GEO strategy.
-
Wildcard group decided24 files A named rule stops GPTBot, nothing names ChatGPT-User, and the general group permits the root. -
Its own group decided12 files A group naming ChatGPT-User allows it. Written down rather than inferred, on figma.com, svt.se, ebay.com and nine others. -
No group applied2 files No group matches ChatGPT-User, which the standard treats as allowed. -
The reverse case2 files GPTBot allowed at the root and ChatGPT-User refused, on cambridge.org and gmarket.co.kr.
Three vendors ship a user-initiated fetcher and document it three ways
OpenAI is not alone in this. Of the fifteen tokens this scanner evaluates, three belong to a fetcher whose job is to retrieve a page because a person asked: ChatGPT-User, Claude-User and Perplexity-User. All three vendors publish a page describing it. No two of those pages say the same thing about whether the robots exclusion standard binds it, and reading all three is the only way to know what a Disallow line is worth in each case.
Anthropic's crawler page, dated 7 April 2026 and read at source on 28 September 2026, is the unambiguous one. It lists three bots, says of the middle one that Claude-User allows site owners to control which sites can be accessed through these user-initiated requests, and states as a principle that Anthropic's bots respect do not crawl signals by honoring industry standard directives in robots.txt. It also spells out the cost of using that control, saying that disabling Claude-User prevents the system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search. That is a vendor documenting a trade rather than selling a setting, and it is the clearest of the three.
Perplexity's crawler page, read at source the same day, contradicts itself inside one table cell. It says that Perplexity-User controls which sites these user requests can access, and then, four sentences later in the same cell, that since a user requested the fetch, this fetcher generally ignores robots.txt rules. Both sentences are on the page. A site owner reading the first and skipping the second would write a rule expecting it to hold. The page is more useful than that reading suggests, because it also publishes IP ranges and gives firewall configuration steps for Cloudflare and AWS, which is a control that does not depend on the fetcher choosing to honour anything.
OpenAI sits between the two: robots.txt rules may not apply, which is weaker than Anthropic's commitment and stronger than Perplexity's disclaimer. For a site owner the practical consequence is that there is no single sentence to follow. The rule you write for Claude-User is documented to bind. The identical rule written for Perplexity-User is documented not to. The one written for ChatGPT-User may or may not, and the vendor has not said which. Our per-platform pages for Claude and Perplexity carry the current wording for each, because a citation strategy built on one vendor's policy generalises to no other, and that is as true of answer engine optimisation advice as it is of blocking advice.
The corpus reflects none of this nuance. 103 files name ChatGPT-User, 48 name Perplexity-User and 41 name Claude-User, and 29 files name all three. Where a site does address the class, it tends to address it uniformly, which is a reasonable response to three vendors whose documentation cannot be reconciled.
| Token | Vendor | What the vendor's page says about robots.txt | Files naming it | Files refusing it at the root |
|---|---|---|---|---|
| ChatGPT-User | OpenAI | Robots.txt rules may not apply | 103 | 45 |
| Claude-User | Anthropic | Bots respect do not crawl signals by honoring directives in robots.txt | 41 | 22 |
| Perplexity-User | Perplexity | This fetcher generally ignores robots.txt rules | 48 | 25 |
What a Disallow line cannot do, and what this run did not measure
The honest summary of the measurement is narrow. This run read 1,025 published files and worked out what each one permits for four named tokens on 28 September 2026. It sent no request as GPTBot, as ChatGPT-User or as anything other than this scanner's own declared token, which announces itself and is described on our bot page. It observed no OpenAI fetch, holds no server logs, and therefore supports no claim at all about whether any vendor's software obeys any of these files. Anybody arguing from these numbers to a conclusion about compliance is arguing past the evidence.
What the numbers do support is a statement about configuration, and it is unflattering. On 38 files a named rule stops the training crawler and the agent that retrieves pages for answers is allowed through the front door, and on 26 of those 38 nothing in the file suggests anybody intended that. The most common single cause is the least interesting one: a site copied a token list, the list predated the user-initiated class, and the wildcard group quietly decided the rest. That is a maintenance problem rather than a policy dispute, and it is visible in about four percent of the files here.
There is a control that does not rest on a vendor's goodwill, and OpenAI publishes the raw material for it. The documentation links a JSON list of IP prefixes for each agent, and on 28 September 2026 the list at openai.com/chatgpt-user.json carried 230 prefixes and a creation time of 25 September 2026, against 18 prefixes and 22 September 2026 for the GPTBot list. Filtering or allowing by address is a different kind of statement from a Disallow line, because it does not ask the client to cooperate. We counted the same four lists in more detail when we opened and counted OpenAI's crawler IP files, and the standing caveat is the one to keep: those lists change without notice, so a firewall rule built from them is a subscription rather than a setting.
Two further limits belong on this measurement. First, a user-agent string is a claim and not an identity, which we have written about at length in a user agent is a claim, not an identity, so a rule keyed to a token is only as good as the requester's honesty about its own name. Second, robots.txt is not the only thing between a fetcher and a page: 103 of 1,089 sites in an earlier sweep served an unknown bot and then refused GPTBot at the infrastructure layer, and on 76 of those the site's own file permitted it. Where a CDN and a file disagree, the CDN wins and the file is documentation.
For anyone checking their own site today, the useful exercise is small. Read your file and count the tokens in it. If it names GPTBot and not ChatGPT-User, you have opted out of training and not out of retrieval, and you should decide which of those you meant. If it names neither, the wildcard group is answering for you. And if you want to know what a specific agent actually gets, ask the site as that agent rather than reasoning from the file, which is the difference between the two numbers we reported when the answer changed with the agent on 4 of 29 sites. Reading permission is cheap and it is not the same thing as measuring AI visibility.
- Disallow line naming ChatGPT-User Rests on the vendor honouring it. OpenAI's page says robots.txt rules may not apply to this agent. 103 of 1,025 corpus files carry one.
- Disallow line naming GPTBot only Addresses training. On 38 corpus files it stops GPTBot at the root and leaves ChatGPT-User allowed there.
- Wildcard group with no AI token named Decided the ChatGPT-User verdict on 24 of those 38 files. It permits whatever it does not refuse.
- Filtering by published IP prefix openai.com/chatgpt-user.json listed 230 prefixes on 28 September 2026, created 25 September 2026. Does not ask the client to cooperate, and changes without notice.
- Infrastructure or CDN rule on the user agent Applies before the file is consulted. Only as reliable as the user-agent string, which is a claim rather than an identity.
Lantad
Published .
The instruction for keeping a page out of ChatGPT is always the same one line. Name the bot, write Disallow, save the file. It is the first thing a hosting help centre says and the first thing our own robots.txt tester will evaluate for you. What that instruction leaves out is that ChatGPT is not one client, and the file does not govern all of them equally.
Common questions
Can I block ChatGPT with robots.txt?
You can block three of the four agents OpenAI documents, and the fourth carries a caveat from OpenAI itself. Its crawler documentation, read at source on 28 September 2026, says that disallowing GPTBot signals that content should not be used in training and recommends using OAI-SearchBot to manage search opt-outs, but says of ChatGPT-User that because these actions are initiated by a user, robots.txt rules may not apply. So a Disallow line for GPTBot is a training opt-out, a Disallow line for OAI-SearchBot is a search opt-out, and a Disallow line for ChatGPT-User is a request whose standing the vendor has not confirmed.
What is the difference between GPTBot and ChatGPT-User?
GPTBot crawls the web automatically to collect training data for foundation models. ChatGPT-User fetches a single page because somebody in a chat asked a question that needed it, and OpenAI's page states that it is not used for crawling the web in an automatic fashion. They are separate tokens with separate published IP lists, and a rule for one has no effect on the other. In this corpus 77 of the 172 files naming GPTBot named ChatGPT-User nowhere, so that separation is where most of the misconfiguration sits.
Do Claude and Perplexity treat user-initiated fetches the same way?
No, and their own pages say so in opposite directions. Anthropic's page, dated 7 April 2026, says its bots respect do not crawl signals by honoring industry standard directives in robots.txt and that Claude-User lets site owners control which sites these user-initiated requests can reach. Perplexity's crawler page says that Perplexity-User controls which sites these user requests can access and also, in the same table cell, that since a user requested the fetch, this fetcher generally ignores robots.txt rules. One rule written three times has three different documented meanings.
Did Lantad measure whether ChatGPT obeys these files?
No. This run requested each site's robots.txt once as LantadBot and evaluated paths against it with the scanner's own parser. No request was sent as any OpenAI agent, no OpenAI fetch was observed and no server logs were held, so nothing here supports a claim about compliance. The figures describe what 1,025 files permitted on 28 September 2026. Separating published permission from observed behaviour is deliberate, because the two have been shown to differ and only one of them is cheap to read.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.