BlogFindings
OAI-SearchBot: 16 of 140 robots.txt files named it, and 25 named GPTBot
Lantad requested /robots.txt once from each of 170 hostnames on 8 September 2026. 140 answered with a file that parsed as text. 25 of those name GPTBot, 16 name OAI-SearchBot, and 11 name GPTBot without ever naming the crawler OpenAI's documentation ties to appearing in ChatGPT search answers.
Two kinds of claim sit in this post and they are kept apart on purpose. The counts are ours, measured on one day from one network location, and the method and its limits have a section to themselves at the end. Everything about what each token governs is read from OpenAI's own crawler documentation, opened the same day: we have not observed any OpenAI crawler arriving at any site, and nothing here measures what ChatGPT does with a page. The reason the two halves belong together is that the documentation draws a distinction which most of the files, on the evidence of this sample, do not.
In short
- Lantad requested /robots.txt once from each of 170 hostnames on 8 September 2026, and across the 140 files that answered HTTP 200 and parsed as text, GPTBot is named in 25, ChatGPT-User in 21, OAI-SearchBot in 16 and OAI-AdsBot in 2.
- OAI-SearchBot is the token OpenAI's crawler documentation, read on 8 September 2026, ties to search visibility: the page states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.
- Eleven of the 140 files name GPTBot and never name OAI-SearchBot, and five of those eleven disallow the whole site to GPTBot while leaving the search crawler to whatever the wildcard group says, which in all five cases is no root disallow.
- OpenAI's documentation states that ChatGPT-User is not used to determine whether content may appear in Search, and ChatGPT-User is named in 21 of the 140 files against 16 for OAI-SearchBot.
- OpenAI's documentation states that if a site has allowed both GPTBot and OAI-SearchBot it may use the results from just one crawl for both use cases to avoid duplicative crawling, so a count of requests per token in an access log is not a count of purposes.
| OpenAI token | What OpenAI's page ties it to | Files naming it | Of those, root disallowed |
|---|---|---|---|
| GPTBot | Training of its foundation models | 25 | 11 |
| ChatGPT-User | Fetches a user's question triggers | 21 | 6 |
| OAI-SearchBot | Appearing in ChatGPT search answers | 16 | 3 |
| OAI-AdsBot | Safety checks on submitted ad pages | 2 | 1 |
What is OAI-SearchBot, and what does allowing it decide?
OpenAI documents four separate tokens on one page, and that page carries no last updated date anywhere in its text, so 8 September 2026 is the only date we can attach to what it said. OpenAI's crawler documentation describes GPTBot as used to crawl content that may be used in training our generative AI foundation models, and states that disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models. OAI-SearchBot is described in different terms entirely: it is for search, it is used to surface websites in search results in ChatGPT's search features, and sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.
That last clause is worth keeping in view. Opting the search crawler out does not remove a site from ChatGPT altogether, because the documentation reserves the possibility of a navigational link. What it removes is the answer, and the answer is the surface most of the AI visibility conversation is actually about.
The page then recommends a configuration in plain terms: to help ensure a site appears in search results, allow OAI-SearchBot in robots.txt and allow requests from the published IP ranges. It also names the split most site owners say they want, stating that a site owner can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI's generative AI foundation models. That sentence is the reason the two tokens exist separately, and the arrangement is the one our directory of AI crawlers exists to make legible. One more sentence sets a clock on any change a reader makes today: for search results, OpenAI says it can take about 24 hours from a robots.txt update for its systems to adjust. A rule in that file is a request to a program that reads it on its own schedule, which is true of every AI crawler and not a peculiarity of this one.
| Token | In OpenAI's words | What the page says about robots.txt |
|---|---|---|
| GPTBot | Used to crawl content that may be used in training our generative AI foundation models | Disallowing it indicates content should not be used in training |
| OAI-SearchBot | Used to surface websites in search results in ChatGPT's search features | Allowing it is recommended for appearing in search results |
| ChatGPT-User | Used for certain user actions in ChatGPT and Custom GPTs | Because the action is user initiated, rules may not apply |
| OAI-AdsBot | Used to validate the safety of web pages submitted as ads on ChatGPT | Not addressed in the entry |
What 170 robots.txt files said about OpenAI's four tokens
Method first, because the counts mean little without it. On 8 September 2026 Lantad requested one robots.txt per hostname over HTTPS for 170 hostnames, sending LantadBot/1.0 as the user agent, following redirects, with a 25 second timeout on each request. The 170 were chosen by hand to spread across retail, publishing, software documentation, universities, government, banking, travel and the standards bodies. They are not a random sample and no ranked list was used, so the set describes large well known sites and nothing wider than that.
141 hostnames answered HTTP 200. One of those, www.terraform.io, returned an HTML document rather than a robots file, which leaves 140 files that parsed as text. Of the remainder, 17 answered 403, though two of those came from a network policy between our runner and the host rather than from the site itself, which we can tell because the response body said so. Five answered 503, four answered 404, one answered 444, and two hostnames failed at the network layer, one refusing the connection and one timing out at 25 seconds. The difference between a 404 and a 503 on this file is not cosmetic, and we have written before about how the two statuses pull a crawler in opposite directions.
Each file was split into groups the way RFC 9309 describes: one or more consecutive user-agent lines open a group, the rules that follow belong to every agent named in it, and token matching is case insensitive. The 140 files hold 1,311 groups between them. 138 carry a wildcard group and 6 disallow the site root inside it. The median file is 1,484 bytes and the largest is 383,108. One file, www.visa.com, answered 200 with an empty body.
Then the counts. 27 of the 140 name at least one of OpenAI's four tokens, so 113 name none of them. 105 of the 140 name none of the 16 tokens we looked for at all, which is the 15 this scanner evaluates plus OAI-AdsBot. That proportion is close to what we found in a different sample of WordPress sites, where 16 of 24 named no AI crawler, and it is the ordinary case: most sites have never written a rule about an AI crawler at all.
Eleven files named GPTBot and never named the search crawler
Eleven of the 140 files name GPTBot and never name OAI-SearchBot anywhere in the file. Five of those eleven disallow the whole site to GPTBot: medium.com, www.canva.com, www.goodreads.com, www.nature.com and www.science.org. The medium.com group is compact enough to read in one line, naming GPTBot and meta-externalagent together above Disallow: / and Allow: /about. www.nature.com is the sharpest case in the set. It names eleven of the sixteen tokens we looked for, refuses GPTBot and ChatGPT-User and Google-Extended the whole site, and never mentions OAI-SearchBot. Its wildcard group does not disallow the root.
What follows from that is not a criticism of any of those files, because we do not know what their authors intended. It is a statement about effect. A crawler named in no group falls to the wildcard group, which is the group most AI crawlers end up reading, and in all five of these files the wildcard group leaves the site root open. On OpenAI's own account, a site that has not opted OAI-SearchBot out remains eligible to be shown in ChatGPT search answers. So a publisher who has written rules against eleven AI crawlers to keep its content out of models has, on the reading of its own file, left the ChatGPT search crawler running.
Whether that is what each of them wanted is a question only they can answer, and there is a coherent position behind it: refuse the training crawl, keep the citation. Two of the five sit in scientific publishing, where that position is common and defensible, and it is the same position OpenAI's page describes as available.
The narrower point is about legibility. If a file names GPTBot, ClaudeBot, PerplexityBot, Bytespider, CCBot and Amazonbot and stops there, no reader can tell whether OAI-SearchBot was left out deliberately or was never on the list somebody worked from. A rule you did not write looks identical to a rule you decided against, which is the same class of problem as a token spelled so that no crawler can match it or a pattern whose trailing character changes nothing. The file is the only artefact anybody else can read, so anything true only in the author's head is not in it.
| Site | Tokens named | GPTBot gets | OAI-SearchBot | Wildcard group |
|---|---|---|---|---|
| www.nature.com | 11 | Disallow: / | Not named | No root disallow |
| www.canva.com | 9 | Disallow: / | Not named | No root disallow |
| medium.com | 6 | Disallow: / with Allow: /about | Not named | No root disallow |
| www.science.org | 6 | Disallow: / | Not named | No root disallow |
| www.goodreads.com | 2 | Disallow: / | Not named | No root disallow |
ChatGPT-User is named more often than the token OpenAI ties to search
ChatGPT-User appears in 21 of the 140 files, five more than OAI-SearchBot, and six of those 21 disallow it the site root. Five files name ChatGPT-User without naming OAI-SearchBot at all: redis.io, www.cloudflare.com, www.intel.com, www.nature.com and www.science.org. Two of those five, nature.com and science.org, disallow the root to ChatGPT-User.
OpenAI's page is explicit about what a rule for that token does and does not achieve, in two sentences sitting next to each other. The first is that because these actions are initiated by a user, robots.txt rules may not apply to ChatGPT-User. The second is more directly useful and almost never quoted: ChatGPT-User is not used to determine whether content may appear in Search, and the page asks readers to use OAI-SearchBot in robots.txt for managing Search opt outs and automatic crawl.
Read together, those two sentences say that a rule written for ChatGPT-User aims at the fetch a person's question triggers, not at the index behind ChatGPT's search answers, and that OpenAI does not undertake to honour it. A site that disallowed ChatGPT-User and never mentioned OAI-SearchBot has written the rule the vendor binds itself to least and skipped the one it binds itself to most. The published address space behind these bots is lopsided in the same direction: when we read OpenAI's four IP range files in July, ChatGPT-User carried 286 prefixes against 21 for GPTBot.
None of this turns a naming count into a verdict, and it should not be read as one. A file that names a token has not necessarily blocked it, and a file that omits one has not necessarily allowed everything, because the wildcard group still applies. That is why the per agent verdicts are a separate exercise, and when we ran them across a smaller sample of sites the three OpenAI agents received the same answer on 25 of 29 files and different answers on 4. What the counts here measure is attention: which tokens site owners have thought about hard enough to type. On that measure the crawler OpenAI ties to search visibility is the one least often typed, and it is worth remembering throughout that a token in a file is a claim about intent rather than an identity check.
-
OAI-SearchBotNamed in 16 of 140 The token OpenAI ties to being shown in ChatGPT search answers, and the least often named of the three main ones. -
GPTBotNamed in 25 of 140 Training. Eleven of the 25 disallow it the site root, the strongest signal of deliberate authorship in the set. -
ChatGPT-UserNamed in 21 of 140 OpenAI states robots.txt rules may not apply to it and that it does not determine whether content may appear in Search. -
OAI-AdsBotNamed in 2 of 140 Only www.ebay.com and www.expedia.com wrote a rule for it, and only the first disallows the root.
One crawl for both use cases: what OpenAI says when both bots are allowed
The sentence that closes the GPTBot entry is the one nobody quotes, and it changes what an access log can tell you. Having described the arrangement where a site allows the search crawler and disallows the training crawler, the same documentation page says that if your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling.
Both halves of that are load bearing. It applies only where both bots are allowed. It says may rather than will. And it describes one fetch being reused for two purposes, which is a statement about efficiency rather than a promise about anything else. What the sentence does not say is anything at all about the reverse case: a site that disallows GPTBot and allows OAI-SearchBot is not told here what happens to the search crawl beyond search. Reading a guarantee into that silence would be the move our method exists to refuse, so we will not: it is a question for OpenAI, and the honest report is that the page does not answer it.
The measured half fits alongside it. Fourteen of the 140 files name both tokens, and in ten of those fourteen the group naming GPTBot and the group naming OAI-SearchBot carry the same number of rules and the same treatment of the site root. Most sites that have thought about both have therefore decided the same thing about both, which is precisely the case where OpenAI says one crawl may serve two purposes. Only four files treat the two differently, and three of those, www.ebay.com, www.linkedin.com and www.tripadvisor.com, disallow the site root to GPTBot while giving OAI-SearchBot a path level rule set with no root disallow. That is the split OpenAI's documentation describes, and in this sample three sites reached it by naming both tokens while eleven reached something like it by naming one.
For anybody reading server logs, the consequence is that requests counted per token are not purposes counted per token. The tokens that arrive are already a poor guide to the fleet behind them, only 6 of the 15 in our registry publish a complete user agent string that can be matched literally, which we found by reading nine vendors' documentation, and the vendor now states that one fetch may be doing two jobs.
Flow: Your robots.txt to Both tokens allowed; Your robots.txt to GPTBot disallowed; Both tokens allowed (may reuse) to One crawl may serve both; GPTBot disallowed to Search crawl only; One crawl may serve both to ChatGPT search answers; Search crawl only to ChatGPT search answers.
What this measurement does not show, and what your own file would
Four limits, stated plainly, because naming counts are cheap to misread.
We counted names and read the rules inside the groups that name them. We did not compute a per path verdict for every crawler against every file, and we fetched no page other than the robots file itself, so nothing here says whether any of these sites is readable once a crawler is through the door. Readability is a separate measurement and it is what our method covers.
We requested each file once, from one network location, with our own user agent. A site that varies its answer by requester would not show that here, and 15 hostnames returned 403 to a self identifying crawler asking for a public file, which is a reminder that the file is not always reachable in the first place. We have also measured how many different robots files one company can serve across its hostnames, and the answer was not one, so a single request to a single hostname describes that hostname and no more.
A robots.txt is a snapshot. Any of these files can change tomorrow, and OpenAI's page puts roughly 24 hours between an edit and its search systems adjusting, which is one of several reasons an edit does not reach a crawler the moment you save it.
What a reader can do with this in ten minutes is concrete. Open your own robots.txt and search it for all four OpenAI tokens. If OAI-SearchBot is absent, decide whether absence is what you meant, and write the decision down either way so the next person to edit the file inherits it. Run the file through the robots.txt tester to see the verdict each agent actually receives rather than the rule you believe you wrote, then look at what a crawler reads on the page once it is allowed in. If the goal is a citation rather than a crawl, the ChatGPT platform guide sets out what else has to be true, and the same exercise for Anthropic's three tokens is in the Claude post of 7 September.
- Which OpenAI tokens each file names Counted across 140 parseable files: GPTBot 25, ChatGPT-User 21, OAI-SearchBot 16, OAI-AdsBot 2.
- Whether a naming group disallows the site root Read from the rules inside each group that names the token, not inferred from the token.
- Whether a wildcard group exists and blocks the root 138 of 140 carry a wildcard group and 6 disallow the root inside it.
- The HTTP status each hostname returned 141 answered 200, 17 answered 403, 5 answered 503, 4 answered 404, 1 answered 444, 2 failed at the network layer.
- A per path verdict for every crawler Not run here. Naming a token is not blocking it, and omitting one is not allowing everything.
- Anything about the pages behind the robots file No page was fetched, so this says nothing about what a crawler could read once allowed.
- What OpenAI's crawlers actually did Not observable from outside a site's own logs. Every statement about crawler behaviour here is quoted from OpenAI's documentation.
Lantad
Published .
OAI-SearchBot is the OpenAI token that decides whether a site can be shown in ChatGPT search answers, and it is the one most robots.txt files do not mention. On 8 September 2026 Lantad requested /robots.txt once from each of 170 hostnames, sending the scanner's own user agent and following redirects. 141 answered HTTP 200 and 140 of those returned plain text that parsed. Across those 140 files, GPTBot is named 25 times, ChatGPT-User 21 times, OAI-SearchBot 16 times and OAI-AdsBot twice. Eleven files name GPTBot and never name OAI-SearchBot anywhere, and five of those eleven refuse GPTBot the whole site.
Common questions
What is OAI-SearchBot and do I need to allow it?
OAI-SearchBot is the OpenAI crawler for search. Its documentation, read on 8 September 2026, states that it is used to surface websites in search results in ChatGPT's search features and that sites opted out of it will not be shown in ChatGPT search answers, though can still appear as navigational links. OpenAI recommends allowing it in robots.txt, and allowing requests from its published IP ranges, for a site that wants to appear in those results. Whether you want that is your decision, not a technical question.
Does blocking GPTBot also block ChatGPT search?
Not according to OpenAI's documentation. The page states that a site owner can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training. In the 140 files Lantad read on 8 September 2026, three sites had written exactly that split by naming both tokens, and eleven more named GPTBot alone, which leaves OAI-SearchBot to the wildcard group rather than opting it out.
Does a rule for ChatGPT-User keep my pages out of ChatGPT?
OpenAI's page says two things that bear on this. Because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply to it. And ChatGPT-User is not used to determine whether content may appear in Search, with the page directing readers to OAI-SearchBot for managing Search opt outs. In this sample ChatGPT-User was named in 21 of 140 files, more often than OAI-SearchBot at 16.
If I allow both OpenAI crawlers, do I get two crawls?
OpenAI's documentation says not necessarily: if your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling. That sentence applies only where both are allowed and it says may rather than will. The practical consequence for a site owner is that counting requests per token in an access log does not count purposes, because one fetch may be serving two of them.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.