BlogFindings

Does blocking AI training affect search rankings: 52 of 54 sites refusing Google-Extended still allow Googlebot

Cloudflare announced a Disallow AI Training setting and an Accountable designation on 15 September 2026, built on the promise that refusing training costs a site nothing in search. Lantad requested /robots.txt once from each of 1,027 hostnames on 16 September 2026. 700 files parsed into at least one group, 54 of them disallow Google-Extended at the site root, and on 52 of those 54 Googlebot is still allowed.

14 min read Lantad

On 15 September 2026 Cloudflare published a post introducing a setting called Disallow AI Training and, alongside it, a designation it calls Accountable. The argument is that some of the largest crawlers on the web are mixed-use, serving search and training from one fetch, so refusing one has meant refusing both. The designation is Cloudflare's attempt to unpick that by getting operators to commit to four things, one of which is an explicit assurance that refusing training will not move search results. An earlier post here covered the ads-dependent defaults that took effect on the same date. This one asks a narrower question that a scanner can answer from outside: what do real robots.txt files already do about the split, and which parts of the promise can anyone check? On 16 September 2026 we requested /robots.txt over HTTPS, once, from each of the 1,027 hostnames in this repository's committed industry corpus, as LantadBot/1.0 with redirects followed and a twenty second timeout, from one network location.

In short

  • Does blocking AI training affect search rankings: Google's own crawler documentation, carrying Last updated 2026-07-14 UTC, states that Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.
  • Cloudflare published a designation it calls Accountable on 15 September 2026 with four requirements, the fourth of which is an assurance that opting out of AI training will not affect traditional search results. It names Applebot, Bingbot and Googlebot as Accountable mixed-use crawlers.
  • Lantad requested /robots.txt once from each of the 1,027 hostnames in this repository's industry corpus on 16 September 2026. 700 of the responses parsed into at least one group, 54 of those disallow Google-Extended at the site root, and 52 of the 54 leave Googlebot allowed.
  • The two files in the sample that do refuse Googlebot, wa.gov and helsinki.fi, refuse it through a blanket Disallow: / under User-agent: * rather than through any rule naming Google. No file in the 700 writes a rule aimed at Googlebot alone.
  • Of the three mixed-use crawlers Cloudflare designated Accountable, robots.txt carries a training opt-out token for two on 16 September 2026. Cloudflare's post states that Microsoft is building support for a no-training preference in robots.txt and targets early 2027.
OperatorTraining opt-outWhere it livesStated effect on search
Google (Googlebot)Google-Extendedrobots.txt tokenDoes not impact inclusion, not a ranking signal
Apple (Applebot)Applebot-Extendedrobots.txt tokenContent remains discoverable in Spotlight, Siri and Safari
Microsoft (Bingbot)NOARCHIVE meta tagPage HTML, not robots.txtCloudflare reports robots.txt support targeted for early 2027
Amazon, Anthropic, Meta, OpenAISeparate training crawlerrobots.txt token per crawlerCloudflare states blocking training does not affect search
The three mixed-use crawlers Cloudflare designated Accountable on 15 September 2026, with the training opt-out each operator documents. Operator statements read at source on 16 September 2026; the Bing row is from Cloudflare's post, because Bing's own help page could not be read without JavaScript.

Does blocking AI training affect search rankings?

For Google the answer is published, and it is the flattest sentence on the page. Google's list of common crawlers, carrying Last updated 2026-07-14 UTC, says of the Google-Extended product token that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". The same entry explains why the token behaves oddly in a log file: Google-Extended has no HTTP user agent string of its own, crawling is done with existing Google user agent strings, and the robots.txt token is used in a control capacity. It is a preference, not a fetcher. That distinction was the subject of an earlier post on the tokens that never arrive as a request, and it matters here because it means a site cannot verify the promise by watching its own traffic. Nothing arrives to watch.

Apple publishes the equivalent statement on its Applebot page, at support.apple.com/en-us/119829, read on 16 September 2026. It says web publishers can opt out of training by disallowing Applebot-Extended in robots.txt, and then, directly: even if you disallow Applebot-Extended and tag content with the nosnippet meta tag, your site instructions may still allow Applebot to crawl your webpages, and your content will remain discoverable through Spotlight, Siri and Safari. Apple's host is not on this site's registered outbound list, so that URL is written here as plain text rather than linked.

Microsoft is the one that does not fit. Its AI training preference is expressed through the NOARCHIVE meta tag rather than a robots.txt token, which was the subject of an earlier post on an AI opt out that lives in the page rather than the file, and which is also why Bing sits awkwardly inside a designation whose first requirement names robots.txt. So the honest answer to the question in this heading, on 16 September 2026, is that two of the three largest search operators state in their own documentation that refusing training does not move search, one expresses the preference somewhere robots.txt cannot reach, and none of the three publishes a number. An assurance is not a measurement, and treating one as the other is the mistake this blog exists to avoid. What can be measured is what site owners have already decided, which is the rest of this post and a different question from whether they decided correctly. Anyone reading this to plan generative engine optimization work should hold those apart.

  • GPTBot 70 of 700 OpenAI training crawler
  • CCBot 69 of 700 Common Crawl
  • ClaudeBot 62 of 700 Anthropic training crawler
  • Google-Extended 54 of 700 Google training preference
  • Applebot-Extended 47 of 700 Apple training preference
  • Applebot 11 of 700 Apple search crawler
  • Bingbot 5 of 700 Microsoft search crawler
  • Googlebot 2 of 700 Google search crawler
Disallowed at the site root, evaluated at path / for each token against 700 robots.txt files that parsed into at least one group. One GET per hostname, 1,027 hostnames asked, 16 September 2026.

What Cloudflare's Accountable designation asks an operator for

Cloudflare's post sets out four requirements a bot operator must meet or commit to meeting. A mechanism for site owners to opt out of AI training, through robots.txt or a similar standard. A mechanism to opt out of AI summaries, set with the operator directly and, next year, through Cloudflare. URL-level visibility into which pages were made available for training, along with metrics showing how content appeared in search. And an assurance that opting out of AI training will not affect traditional search results. Apple, Google and Microsoft are described as demonstrating that they meet the qualifications, each combining capabilities available today with time-bound commitments for the rest.

Read as a scanner would read it, those four are not the same kind of claim. The first is a published mechanism, so anyone can check it by fetching a documentation page and then testing a file. The other three are statements about what happens inside an operator. URL-level visibility is a dashboard the site owner sees only after signing in, and Cloudflare's own post records that Apple does not provide that tool yet and that Google expects additional URL-level transparency tools in the weeks to come. The summaries opt-out is set with each operator separately today. The assurance about search is a sentence, and the value of a sentence depends entirely on who wrote it. That is not a criticism of the designation, which is more concrete than anything the category had before. It is a limit on what an external scan of a site can confirm, and it belongs in the open rather than in a footnote.

The mechanical part of the announcement is easier to pin down. Disallow AI Training is named for the Disallow directive it publishes in robots.txt, written through Bot Preference Sync, and Cloudflare states that Managed Robots.txt, the feature that generated and maintained the file for customers, is deprecated in favour of it. An earlier post here noted that the token list in a generated robots.txt moves without the site owner touching anything, which is the same property viewed from the site's side, and another covered the Content-signal line that asks while the Disallow lines under it block. A preference written into your file by a vendor is still your file to a crawler. It is worth knowing which lines you wrote.

  • Training opt-out through robots.txt or similar Checkable. The operator documents a token, and any client can fetch the file and evaluate it. This is the only one of the four a scan reaches.
  • Opt-out from AI summaries, set with the operator Not checkable from outside. Set in each operator's own surface today; Cloudflare says a single control through Cloudflare is a goal for next year.
  • URL-level visibility into pages used for training Not checkable from outside. A signed-in report. Cloudflare's post says Apple does not offer the tool yet and Google expects more in the weeks to come.
  • Assurance that opting out will not affect search A statement, not an observation. Google and Apple publish it on their own pages; nobody publishes a figure behind it.
The four requirements Cloudflare published on 15 September 2026, with whether an external scan of a site can confirm each one. The verdicts in the detail column are this post's reading, not Cloudflare's.

What 700 robots.txt files said on 16 September 2026

The frame is this repository's committed industry corpus of 1,027 hostnames across eight sectors, which is an editorial sampling frame weighted toward large organisations rather than a random draw from the web. Each hostname was asked once for /robots.txt over HTTPS as LantadBot/1.0, the user agent described on our crawler conduct page, with redirects followed and a twenty second timeout. 28 requests failed at the transport layer. 741 answered HTTP 200, 171 answered 403, 46 answered 503 and 25 answered 404. Of the 741, 32 returned HTML or XML rather than a robots file, leaving 709 textual responses of which 700 parsed into at least one group. Every count below is out of those 700, and each was evaluated at the path / using this repository's own parser, which follows RFC 9309 with the Google-documented refinements on group selection.

Fifty four of the 700 disallow Google-Extended at the site root. On 52 of those 54, Googlebot is still allowed. The two exceptions are wa.gov and helsinki.fi, and neither is a decision about Google: in both files the rule that catches Googlebot is a bare Disallow: / on line two under User-agent: *, which catches everything else as well. Not one file in the 700 writes a rule aimed at Googlebot on its own. Read the other way round, the split that Cloudflare's new setting exists to make easy is already what almost everyone who has made the choice has chosen.

The same holds for the other two Accountable search crawlers, with more noise. Twelve of the 700 disallow at least one of Googlebot, Bingbot or Applebot at the root, which is 1.7 percent. Cloudflare's post states that less than 1 percent of Cloudflare sites choose to block Search bots. Those are not the same measurement and should not be reported as agreeing: Cloudflare is counting configurations across its own network, and this is counting robots.txt files in a corpus chosen for a different purpose. The direction is the same and the figures are not comparable. The training side lands similarly: 95 of the 700, or 13.6 percent, disallow at least one of the nine training-side tokens this scanner evaluates, against Cloudflare's figure of 17 percent of sites enabling some mechanism to block training, where a mechanism includes settings that never appear in a file. Which sites do it is the least surprising part and was covered here before, in the finding that the sites blocking AI crawlers are the ones with editors: news accounts for 41 of the 95 from 63 parsed files, while travel and finance contribute two each from 166 between them. The ranking of tokens has moved little since GPTBot led 82 of 718 files earlier this month.

52 refuse training, keep search

  • Google-Extended disallowed at /
  • Googlebot allowed at /
  • The split Cloudflare's new setting automates
  • Already the choice on 52 of 54 files

2 refuse both

  • wa.gov and helsinki.fi
  • Caught by Disallow: / under User-agent: *
  • No rule in either file names Google
  • Blanket refusal, not a training decision
The 54 files in the 700 that disallow Google-Extended at the site root, split by what the same file does to Googlebot. Measured 16 September 2026.

Two of the three Accountable crawlers have a robots.txt token today

Google-Extended and Applebot-Extended are real product tokens, and the corpus shows both in use: 83 files name Google-Extended and 52 name Applebot-Extended, with 54 and 47 respectively disallowing the root. Sixty three files disallow at least one of the two, 38 disallow both, 16 take only Google's and 9 take only Apple's. Bingbot is named by 47 files, but there is no Bing training token to write, so none of those 47 is a training preference. Cloudflare's post says Microsoft is building the mechanism to respect a no-training preference in robots.txt at the domain level, targeted for early 2027, and states plainly that until that support launches, selecting Disallow AI Training will not automatically convey a no-training preference to Bing through robots.txt. A setting named after a Disallow line does not reach one of the three crawlers the same post designates Accountable, and the sync that publishes those lines can only write what an operator has agreed to read.

The two tokens that do work are not symmetric either, and the asymmetry is in the names. Google chose a stem, Google-Extended, that no other Google token is a prefix of. Apple chose Applebot-Extended, which the shorter token Applebot is a prefix of. That matters because of how group selection works: Google's robots.txt specification page, carrying Last updated 2026-08-31 UTC, states that its crawlers find the group with the most specific user agent that matches, and this repository's parser implements the same rule. So a file that disallows Applebot and never mentions Applebot-Extended still refuses Apple's training preference by inheritance. In this run that happened on exactly one file, snu.ac.kr, so the effect is real and rare rather than widespread. Six files disallow Google-Extended without naming it, all of them through a wildcard group.

The Bing case has one more wrinkle worth recording, because it was measured rather than read. On 16 September 2026 the Bing Webmaster Tools page documenting the robots meta tags Bing supports answered a request from LantadBot with HTTP 200 and 125,562 bytes. With no JavaScript executed, the visible text of that response is 88 characters and 15 words, reading "Bing Webmaster Tools - Help Documentation You need to enable JavaScript to run this app." We therefore did not read Bing's own description of NOARCHIVE and do not quote it here; everything above about Microsoft's position comes from Cloudflare's post. An earlier piece on what Bing does name for grounding is the closest this blog has come to the primary text.

Where a site owner's refusal of AI training lands for each of the three mixed-use crawlers Cloudflare designated Accountable, as documented on 16 September 2026.

What to check on your own site

Three things are worth confirming in your own file, and all three are cheap. First, whether the training tokens you think are in there are the ones actually written, which is what the robots.txt tester evaluates per crawler rather than by eye. Second, whether a wildcard group is doing work you did not intend: six files in this run refuse Google-Extended without ever naming it, and two refuse Googlebot the same way, which is the sort of rule that reads as a training decision in a report and was never one. Third, which of the fifteen tokens in this scanner's crawler registry your file actually addresses, since a file naming five of them is silent about ten.

The inconvenient part goes at the end, because it is the part a vendor post will not carry. The assurance Cloudflare asks operators for is about traditional search results, and that is a narrower promise than it sounds. A SIGIR 2026 study covered here found that AI Overviews retrieved less from sites blocking Google-Extended than ordinary Google Search did, which is entirely consistent with Google's statement, because AI Overviews is not traditional search ranking. Refusing training may leave your blue links exactly where they were and still change how often a generative surface reaches for your page. Nothing in this post measures that, and nothing measured here says whether any crawler obeyed any of these 700 files.

There is also a plainer limitation. Rules written at the site root are the only ones counted above, and an earlier run found that 89 of 145 files naming an AI crawler ruled only on the whole site, so a path-level preference would not show here at all. One file per host, asked once, on one day, from one network location: a transient 403 reads as a permanent refusal, and 171 hosts answered 403 to a request for a public file that RFC 9309 expects to be readable. Those hosts are absent from every figure in this post.

  • What site owners chose Measured 54 of 700 files disallow Google-Extended at the root and 52 of those still allow Googlebot.
  • Whether crawlers obeyed Not measured No access log was read and no crawler was observed fetching any of these hostnames.
  • Whether refusing training costs rankings Not measured Google and Apple state it does not. No figure from anyone, including this run, tests the statement.
  • Whether these sites use Cloudflare Not measured The corpus is chosen by sector, not by provider, so no figure here describes Cloudflare's network.
What this run did and did not establish. Measured 16 September 2026 across 1,027 hostnames in the committed industry corpus.

Written by

Lantad

Published .

The question a site owner actually asks before touching robots.txt is not which token to write. It is whether writing it costs anything. Does blocking AI training affect search rankings, and if nobody can answer that, the safe move is to write nothing and hope. That hesitation is what an AI crawler control has always run into: the mechanism is a single Disallow line, and the consequence of the line has been a matter of trust.

Common questions

Does blocking AI training affect search rankings?

Google states that it does not. Its list of common crawlers, carrying Last updated 2026-07-14 UTC, says Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. Apple's Applebot page, read on 16 September 2026, states that content remains discoverable through Spotlight, Siri and Safari after Applebot-Extended is disallowed. Both are statements by the operator rather than measurements, and neither publishes a figure behind it. Lantad has run no experiment on ranking outcomes and makes no claim about them.

What is Cloudflare's Accountable designation?

A label Cloudflare published on 15 September 2026 for bot operators that meet or commit to meeting four requirements: a training opt-out through robots.txt or a similar standard, an opt-out from AI summaries, URL-level visibility into which pages were made available for training with metrics on how content appeared in search, and an assurance that opting out of training will not affect traditional search results. Cloudflare names Applebot, Bingbot and Googlebot as Accountable mixed-use crawlers, and also categorises the relevant crawlers from Amazon, Anthropic, Meta and OpenAI as Accountable because those operators run separate search and training crawlers.

Can robots.txt tell Bing not to train on my content?

Not on 16 September 2026. Microsoft's AI training preference is expressed through the NOARCHIVE meta tag in a page rather than through a robots.txt token. Cloudflare's post of 15 September 2026 states that Microsoft is building the mechanism to respect a no-training preference in robots.txt at the domain level, targeted for early 2027, and that until then selecting Disallow AI Training will not automatically convey that preference to Bing through robots.txt. Of the 700 robots.txt files Lantad parsed that day, 47 name Bingbot, and none of those 47 can be a training preference because there is no token to write.

How many sites already refuse AI training while keeping search?

In this sample, 54 of 700 parsed robots.txt files disallow Google-Extended at the site root on 16 September 2026, and 52 of those 54 still allow Googlebot. Sixty three files disallow Google-Extended or Applebot-Extended, and 38 disallow both. The frame is this repository's industry corpus of 1,027 hostnames, weighted toward large organisations, so the proportions describe that frame and not the web. A further 171 hostnames answered HTTP 403 to the request and appear in no figure.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.