BlogFindings

Reflectionbot: 1 of 1,068 robots.txt files named it, and 72 of the 77 blocking GPTBot let it in

Reflection AI published crawler documentation for a new training bot in late September 2026. Lantad requested the robots.txt of all 1,419 hostnames in this repository's two committed corpus seed files on 3 October 2026 as LantadBot, and 1,068 returned a parseable file. One named Reflectionbot. On 1,010 of the rest the wildcard group decided, and of the 77 files that refuse GPTBot at the site root, 72 allow Reflectionbot there.

15 min read Lantad

Its page at reflection.ai/bot, read on 3 October 2026, states that Reflectionbot is the web crawler operated by Reflection AI, that it visits publicly available web pages to help build and improve Reflection AI's models, and that Reflectionbot is the only user agent Reflection AI crawls with today. The page asks site owners to express preferences through robots.txt rather than by returning errors such as HTTP 403, and says a robots.txt change is typically picked up within about a day. Reflection AI also publishes its crawler IP ranges as JSON at reflection.ai/bot/ips.json, which on 3 October 2026 carried a creationTime of 22 September 2026 and two IPv4 prefixes, 16.216.94.0/23 and 16.216.88.0/23, and no IPv6 prefixes at all. Those are Reflection AI's statements about its own crawler, reported here rather than measured by us.

What Lantad can measure is the other side: what the files say back. On 3 October 2026 we requested https://host/robots.txt once for each of the 1,419 hostnames in this repository's two committed corpus seed files as LantadBot, following redirects, and evaluated each file against the Reflectionbot token by parsing it into groups and applying the selection rules RFC 9309 sets out, which is the same procedure the robots.txt tester applies to one file at a time. 1,109 answered HTTP 200, 41 of those served HTML where a text file was expected, and 1,068 returned a parseable robots.txt. Every count below is over those 1,068 files on that one date.

In short

  • Reflectionbot is the single crawler token Reflection AI documents, and its own page read on 3 October 2026 describes it as collecting pages to help build and improve Reflection AI's models. Of 1,068 parseable robots.txt files Lantad read that day, exactly 1 named it: slate.com, which refuses it.
  • A token nobody has named is not a token nobody has ruled on. On 1,010 of the 1,068 files the wildcard group decided Reflectionbot's access, carrying a median of 9 rules written for somebody else, and on 57 files no group matched at all, which RFC 9309 defines as no rules applying.
  • 77 of the 1,068 files disallow GPTBot at the site root. 72 of those 77 allow Reflectionbot at the same path, 69 of them through a wildcard group that never names either crawler. 38 of the 72 are news sites, the stratum that blocks most deliberately.
  • The same gap held for every training token in Lantad's registry on 3 October 2026: 68 files blocked ClaudeBot at the root and 64 of those allowed Reflectionbot, 75 blocked CCBot and 70 allowed it, 56 blocked Google-Extended and 52 allowed it.
  • Allowing Reflectionbot buys no citation surface, and that is the part worth stating plainly. OpenAI documents OAI-SearchBot and Anthropic documents Claude-SearchBot as separate search crawlers tied to appearing in answers. Reflection AI's page names one token and calls it the only user agent it crawls with today, so there is no search-side token to permit.
What was checkedWhat it said on 3 October 2026
Product tokenReflectionbot, described as the only user agent Reflection AI crawls with today
Stated purposeVisiting publicly available pages to help build and improve Reflection AI's models
Separate search or citation tokenNone documented
Verification method offeredReverse DNS to hostnames under reflection.ai, plus published IP ranges
IP ranges published2 IPv4 prefixes, 16.216.94.0/23 and 16.216.88.0/23, and no IPv6
creationTime in ips.json2026-09-22T10:00:00.000000
Files naming the token, of 1,0681
What Reflection AI's own documentation at reflection.ai/bot and reflection.ai/bot/ips.json stated when Lantad read both on 3 October 2026, set against what the 1,068 parseable robots.txt files in this repository's corpus said about the token that day. The left column is Reflection AI reporting on itself; only the final row is a Lantad measurement.

What is Reflectionbot, and what does Reflection AI say it does?

Reflection AI's documentation is short and unusually specific about matching, which is worth repeating because it is the part that decides whether a rule works. The page states that Reflectionbot announces itself with a single user-agent token, Reflectionbot, and gives an example request string of the form Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Reflectionbot/1.0; +https://reflection.ai/bot) Chrome/151.0.0.0 Safari/537.36, noting that version numbers may change. It then instructs readers to match on the Reflectionbot token rather than on the full user-agent string, whether in robots.txt or in server logs.

That instruction is correct and it is the one most often got wrong. A full user-agent string carries a browser version that moves, so a rule pinned to the whole string stops matching the first time the vendor bumps a number. This is not a hypothetical failure in this corpus: we have already found 1,004 files naming 2,209 tokens of which 120 named a token Anthropic no longer documents, which is what copied blocklists look like once the vendor has moved on.

The page also offers reverse DNS verification, stating that Reflectionbot requests originate from IP addresses resolving to hostnames under reflection.ai, and publishes the ranges at reflection.ai/bot/ips.json. Reflection AI's host is not registered in this site's outbound link policy, so both URLs appear here as plain text rather than as links, and a reader who wants to confirm any of this should open them directly. When we read that file on 3 October 2026 it held two IPv4 prefixes and nothing else, which is 1,024 addresses across the two /23 blocks and no IPv6 coverage. Its creationTime field read 22 September 2026 while the Last-Modified header on the same response read 2 October 2026, a ten day gap of exactly the kind we catalogued when crawler IP range files carried two dates that disagree.

None of that is a measurement of Reflectionbot's behaviour and this post does not make one. We hold no server logs for other people's sites, we observed no Reflectionbot request, and we did not test whether it honours anything. A crawler's documentation is a claim about intent, and a user agent is a claim rather than an identity in the first place. What can be checked cheaply is whether the vendor gives you the means to verify a request, and here it does: published ranges plus reverse DNS puts Reflection AI ahead of several larger operators, though behind the two that publish a signing key directory for web bot authentication.

GET reflection.ai/bot and /bot/ips.json, 3 October 2026

  • GET https://reflection.ai/bot 200, redirected to /reflectionbot
  • single documented token Reflectionbot
  • stated purpose build and improve Reflection AI's models
  • matching instruction match the token, not the full UA string
  • GET https://reflection.ai/bot/ips.json 200, 171 bytes
  • creationTime 2026-09-22T10:00:00.000000
  • Last-Modified header Fri, 02 Oct 2026 04:51:07 GMT
  • ipv4Prefix entries 2
  • ipv6Prefix entries 0
The two requests Lantad made to Reflection AI's own documentation on 3 October 2026, as LantadBot, redirects followed. Response values are quoted from those responses.

How many robots.txt files name Reflectionbot?

One. Of the 1,068 parseable files, a single robots.txt carried Reflectionbot on a User-agent line: slate.com, which spells it ReflectionBot with a capital B and attaches Disallow: / to the group it sits in. The spelling difference does not matter, because RFC 9309 requires crawlers to use case-insensitive matching when finding the group that matches a product token, so slate.com's rule is a real refusal rather than a near miss.

It is worth saying what kind of file that is. slate.com's robots.txt carries 427 User-agent lines, a consecutive stack of crawler names closed by a single Disallow: /, which is the shape of a maintained publisher blocklist rather than something a site owner typed once. 131 of the 1,068 files in this corpus name twenty or more tokens, and the largest, museodelprado.es, names 639. One of those long lists had Reflectionbot in it eleven days after the creationTime its IP range file carries. The others did not.

A second file, figma.com, matched a naive search for the string reflection, and it is a useful false positive to record rather than quietly drop. Every hit was a page path such as Allow: /templates/reflections-for-meetings/$, not a user agent. Searching a robots.txt for a vendor's name finds page titles as readily as crawler rules, which is why the counts in this post come from parsing User-agent lines into groups rather than from grep.

How far any of this generalises depends on the frame, and it is worth being exact rather than reassuring. These are the 1,419 hostnames this repository seeds its corpus from, grouped into eight industry strata and ten website-builder platform strata, chosen to span the web rather than drawn at random from it. How that frame is built and what it will not support is written up in the crawlability study. Read every rate here as a measurement of 1,068 named sites on one date.

Set the single naming against the files that have demonstrably thought about this subject. 212 of the 1,068 named at least one of the fifteen AI crawler tokens listed in core/src/bots.ts, this scanner's bot registry. That registry is a setting somebody chose rather than a finding, and the AI crawlers reference lists tokens rather than companies for the reason this post keeps running into: the unit a robots.txt file operates on is the token, and a vendor is not a unit at all. 175 named GPTBot, 163 ClaudeBot, 147 Google-Extended, 142 CCBot and 129 Bytespider. Those are not people who have never heard of AI crawling. They are the population that read something, formed a view and wrote it down, and exactly one of them had Reflectionbot. A naming count is a count of strings somebody typed on some past date, a caveat that GPTBot being named in 4.5 percent of files while the wildcard appeared in 77 made concrete, and it understates nothing here: one is one.

  • GPTBot 175 files OpenAI, training
  • ClaudeBot 163 files Anthropic, training
  • Google-Extended 147 files Google, training
  • CCBot 142 files Common Crawl
  • Bytespider 129 files ByteDance
  • PerplexityBot 106 files Perplexity
  • OAI-SearchBot 82 files OpenAI, search
  • Claude-SearchBot 42 files Anthropic, search
  • Reflectionbot 1 files Reflection AI, training
Crawler tokens named on a User-agent line in the 1,068 parseable robots.txt files Lantad read on 3 October 2026. Naming a token means it heads a User-agent line, whatever the rule underneath says.

What decides access for a crawler nobody has named?

Not nothing, which is the point people miss. RFC 9309 settles this in two sentences, and they are worth reading in the order the specification puts them. If no group matches the product token, a crawler must obey the group with a User-agent line of "*", if one is present. If no group matches and there is no "*" group either, or no groups are present at all, no rules apply.

Applied to Reflectionbot across the 1,068 files, that produces three populations. On 1 file the token was named. On 1,010 the wildcard group decided, carrying a median of 9 rules and in one case, boe.es, 12,312 of them. On 57 files no group applied at all, because the file named other crawlers and never wrote a wildcard group, which means Reflectionbot gets everything those files serve. ryanmulligan.dev is the minimal example: its entire robots.txt is one group, User-agent: GPTBot with Disallow: /, so GPTBot gets nothing and every other crawler in existence gets the whole site.

Evaluated at the site root, 14 of the 1,068 files disallowed Reflectionbot and 1,054 allowed it. That 14 is low and it should be: a wildcard group carrying Disallow: / refuses Googlebot too, so almost nobody writes one. The 14 that did are coralvilleanimalhospital.com, botcity.dev, amsterdam.nl, wa.gov, helsinki.fi, sciencedirect.com, scielo.org, pennmedicine.org, eluniversal.com.mx, slate.com, theregister.com, sap.com, instacart.com and gmarket.co.kr. Thirteen of those refuse it incidentally, as a side effect of refusing everything. One, slate.com, refuses it by name.

Inheriting the wildcard group is not the same as being unrestricted, and it is not the same as being governed either. A median of 9 rules means most of these sites hand Reflectionbot a housekeeping list: a /search path, an account area, a few asset directories. Those rules were written for search engines years before Reflection AI existed, and whatever they happen to cover is now this crawler's access policy. That mechanism, a rule deciding a crawler's fate without ever naming it, is the same one behind 559 of the 581 pages GPTBot lost being closed by a rule that never named it, running here in the permissive direction instead of the restrictive one.

Group selection for Reflectionbot across the 1,068 parseable robots.txt files Lantad read on 3 October 2026, following the order RFC 9309 sets out. Counts are the measured populations; the diagram shows the decision path rather than any one site.

72 of the 77 sites that refuse GPTBot admit Reflectionbot

This is the finding, and it is a finding about mechanism rather than about anybody's carelessness. 175 of the 1,068 files name GPTBot. 77 of those disallow it at the site root, which is a deliberate act: somebody read about OpenAI's training crawler, decided against it and wrote a rule. On 72 of those same 77 files, Reflectionbot is allowed at the root. 69 of the 72 fall to a wildcard group that names neither crawler, and 3 have no applicable group at all.

mxb.dev shows the whole shape in a 25 line file. Its wildcard group disallows /build.txt and /drafts. Underneath it sit separate groups for CCBot, ChatGPT-User, GPTBot, Google-Extended, Omgilibot and FacebookBot, each carrying Disallow: /. Six named crawlers get nothing. Reflectionbot, named nowhere, inherits the wildcard group and reads the entire site except a text file and a drafts directory. Nobody made that decision. The file is doing exactly what it says and exactly what RFC 9309 specifies, and the outcome is the opposite of what the author's six explicit refusals suggest they wanted.

The pattern is not specific to GPTBot. Of the files naming each token, 68 blocked ClaudeBot at the root and 64 of those allowed Reflectionbot, 75 blocked CCBot and 70 allowed it, 68 blocked Bytespider and 63 allowed it, 56 blocked Google-Extended and 52 allowed it, and 41 blocked PerplexityBot and 38 allowed it. Across the 212 files naming any registry token, 148 gave Reflectionbot a different rule set than at least one AI crawler they had named. The asymmetry is the same one we reported when 52 of 54 sites refusing Google-Extended still allowed Googlebot and when 6 of 1,065 files named the Meta token tied to citation: per-token rules do not generalise, because robots.txt has no concept of a category.

The distribution is the part publishers should read twice. 38 of the 72 are news sites. News is the stratum that blocks AI crawlers most deliberately in this corpus, and that is precisely why it is the most exposed: a site with no AI crawler rules at all has nothing to fall out of date, while a site with a carefully maintained list of nine named trainers has a list that is wrong the moment a tenth ships. Education contributed 7 of the 72 and healthcare 7, against 3 from SaaS and 1 from the WordPress stratum. Enumeration costs the people who did the work.

Token named in the fileFiles naming itDisallowed at / for that tokenOf those, Reflectionbot allowed at /
GPTBot1757772
CCBot1427570
ClaudeBot1636864
Bytespider1296863
Google-Extended1475652
PerplexityBot1064138
For each AI crawler token named in the 1,068 parseable robots.txt files Lantad read on 3 October 2026: how many files named it, how many disallowed it at the site root, and how many of those same files allowed Reflectionbot at that path. Verdicts evaluated at the site root under RFC 9309 group selection and longest-match rules.

Allowing Reflectionbot does not buy a citation

There is an argument this blog could make here and will not, because the evidence does not support it. The usual framing for a product that sells AI visibility is that letting a crawler in is how you get mentioned, so a site accidentally blocking one is losing something. For Reflectionbot that argument runs backwards, and saying so is more useful than selling the alarm.

OpenAI's crawler documentation, read on 3 October 2026, separates the two jobs explicitly: it describes OAI-SearchBot as used to surface websites in search results in ChatGPT's search features while disallowing GPTBot indicates that crawled content should not be used for training its generative models. Anthropic's support article does the same, describing ClaudeBot as collecting web content that could contribute to training and Claude-SearchBot as navigating the web to improve search result quality, and states that blocking Claude-SearchBot may reduce a site's visibility in user search results. Those vendors give a site owner two decisions to make, and only one of them is about being cited.

Reflection AI's page offers one token and describes it as collecting pages to help build and improve its models. It states that Reflectionbot is the only user agent Reflection AI crawls with today. So on the documentation as published on 3 October 2026 there is no search-side or citation-side token to allow, which means allowing Reflectionbot is a contribution to a training corpus and nothing else. Whether that is a trade worth making is a judgement about licensing and about what you think training access is worth, not a visibility question, and the guides for getting cited in ChatGPT and in Claude have nothing to say about it because there is nothing yet to say.

That cuts both ways, and the honest version includes the half that is inconvenient for us. A site that wants to refuse training access while keeping citation access has a real decision to make here and the files show almost nobody has made it. But a site that accidentally allows Reflectionbot has not lost a citation, because there is no citation on offer. Access is a necessary condition for being quoted and never a sufficient one, which is the distinction the GEO category collapses most often, and it is why our own scoring methodology treats crawler access as one input rather than as the answer. The practical step is small: evaluate your file once per token rather than once per vendor, which is the ordering problem AI crawler detection turns on, and decide whether your wildcard group says what you would want it to say to a crawler that does not exist yet. On 1,010 of 1,068 files, that group is already the policy.

  • OpenAI Two tokens GPTBot for training, OAI-SearchBot documented as surfacing sites in ChatGPT search features.
  • Anthropic Three tokens ClaudeBot for training, Claude-SearchBot for search quality, Claude-User for user-initiated fetches.
  • Reflection AI One token Reflectionbot, described as the only user agent it crawls with today, for building and improving its models.
  • What allowing it buys No citation surface No search or citation token is documented, so permitting Reflectionbot is training access only.
What each vendor's own documentation published about training and search tokens when Lantad read it on 3 October 2026. This records what the vendors state, not what Lantad measured any crawler doing.

Written by

Lantad

Published .

A blocklist is an enumeration, and an enumeration is only ever as current as the day somebody last edited it. That is the structural weakness of managing AI crawlers by name, and it is usually argued about in the abstract because the moment a new crawler appears and the files have not caught up is hard to catch. Reflection AI published crawler documentation in late September 2026, which makes this one of those moments.

Common questions

What is Reflectionbot?

Reflectionbot is the web crawler operated by Reflection AI. Its documentation at reflection.ai/bot, read on 3 October 2026, states that it visits publicly available web pages to help build and improve Reflection AI's models, that it announces itself with the single user-agent token Reflectionbot, and that this is the only user agent Reflection AI crawls with today. The page instructs site owners to match the Reflectionbot token rather than the full user-agent string, and says Reflectionbot reads and honours robots.txt with changes typically picked up within about a day.

How do I block Reflectionbot in robots.txt?

Write a group that names the token, because robots.txt applies only the most specific matching group. A group headed User-agent: Reflectionbot with Disallow: / refuses it site-wide, and RFC 9309 requires case-insensitive token matching so the capitalisation does not matter. A rule written under a different crawler's name has no effect on it: of the 1,068 parseable robots.txt files Lantad read on 3 October 2026, 77 disallowed GPTBot at the site root and 72 of those allowed Reflectionbot at the same path.

How many sites block Reflectionbot?

In Lantad's sample of 1,068 parseable robots.txt files read on 3 October 2026, 14 disallowed Reflectionbot at the site root and 1,054 allowed it. Only one of the 14 named the token: slate.com. The other 13 refuse it incidentally through a wildcard group carrying Disallow: /, which refuses every crawler including Googlebot. The sample is this repository's committed corpus seed files rather than a random sample of the web, so treat it as a measurement of 1,068 named sites on one date.

Does allowing Reflectionbot help my site get cited by AI?

Not on the documentation as published. Reflection AI names one token and describes it as collecting pages to build and improve its models, with no separate search or citation crawler of the kind OpenAI documents as OAI-SearchBot or Anthropic documents as Claude-SearchBot. On that basis, allowing Reflectionbot grants training access and no citation surface. Lantad has observed no Reflectionbot request and holds no server logs for other sites, so this reports what the vendor publishes rather than what the crawler does.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.