BlogFindings
noarchive: 7 of 1,077 home pages sent the tag two AI vendors read as a training opt out
Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt on 5 October 2026 and read every permitted home page with no JavaScript executed. 1,077 answered HTTP 200 with HTML. 7 of them sent noarchive, the page level directive Amazon's own crawler page reads as do not train and Microsoft's reads as exclusion from Bing Chat answers, and Google's current documentation says it no longer uses at all. None sent nocache.
The directive is noarchive. It was built for a feature that no longer exists, the cached copy of a page a search engine would show beside the result, and Google now says so in plain terms. Amazon and Microsoft both read it as something else. So this run counted it. Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt on 5 October 2026, evaluated each file for this scanner's own token, and read the home page of every host that allowed it with no JavaScript executed and no stylesheet fetched. 1,077 answered HTTP 200 with an HTML content type. Seven of those 1,077 sent noarchive. The number is small, and the interesting part is not the number: it is that three of the seven have already shut the door on the one crawler that would act on it.
In short
- noarchive is a meta robots directive from the era of the search engine cached page, and on 5 October 2026 Lantad found it on 7 of 1,077 readable home pages in this repository's committed corpus. None of the 1,077 sent nocache.
- Amazon's crawler page states that Amazonbot respects a page level noarchive as do not use the page for model training. Microsoft's 22 September 2023 post states that content tagged NOARCHIVE will not be included in Bing Chat answers and will not be used to train its generative AI foundation models.
- Google's robots meta tag documentation, Last updated 2026-03-24 UTC, states that noarchive is no longer used by Google Search because the cached link feature no longer exists, and that nocache is not used by Google Search either.
- Three of the seven sites sending noarchive already disallow Amazonbot outright in robots.txt, so Amazon's crawler never fetches the page that carries the tag. Google's own documentation states these settings can be read and followed only if crawlers are allowed to access the pages that include them.
- Not one of the 1,077 pages addressed any of the fifteen AI crawler tokens in core/src/bots.ts by name in a meta tag, while 180 of the 977 hosts that returned both a readable robots.txt and a readable home page named at least one of them in robots.txt.
| Stage | Hosts | What happened |
|---|---|---|
| Hostnames asked | 1,419 | 392 in ten platform strata, 1,027 in eight industry strata |
| Returned a parseable robots.txt | 1,069 | HTTP 200 with a body that was not HTML; 10 of them empty |
| Disallowed this crawler at the root | 13 | Not fetched further |
| Answered HTTP 403 to the home page | 215 | Refused this crawler |
| Answered HTTP 503 to the home page | 56 | Served no page to this client |
| Never returned a status | 34 | Failed in the client, at DNS, or on the timeout |
| Answered some other status | 23 | 9 of 429, 5 of 202, 4 of 404, 2 of 406 and one each of 401, 451, 498 and 500 |
| Answered 200 with HTML | 1,077 | The denominator for every rate below |
| Sent any crawler directive | 434 | 432 in a meta named robots, 25 in a vendor named meta, 12 in an X-Robots-Tag header |
| Sent noarchive | 7 | 6 news sites and one bank |
| Sent nocache | 0 | Nobody |
What does noarchive do for an AI crawler now?
Three vendors publish an answer and the three answers disagree. Google's robots meta tag documentation, carrying Last updated 2026-03-24 UTC when it was opened on 5 October 2026, is the shortest: the noarchive rule is no longer used by Google Search to control whether a cached link is shown in search results, as the cached link feature no longer exists. The same page says the nocache rule is not used by Google Search. Both directives are, from Google's side, dead letters on a page that still carefully specifies noindex, nosnippet, max-snippet, max-image-preview and the rest.
Microsoft's reading is the opposite and it is older. Microsoft's September 2023 post on controlling content in Bing Chat, dated 22 September 2023, states that content tagged NOARCHIVE will not be included in Bing Chat answers and will not be linked to in the answers, and that Microsoft will not use that content for training its generative AI foundation models. NOCACHE is the softer setting in the same post: content carrying it may be included in Bing Chat answers with only URL, snippet and title displayed, and only URLs, titles and snippets may be used in training. This blog has written about that pair before, in the post on how Microsoft's AI opt out is a meta tag rather than a robots.txt token, and said at the time that most sites would carry neither without measuring whether that was true. This run measures it.
Amazon's own crawler page is the third reading and the most direct about training. It states that Amazonbot respects the link level rel=nofollow directive and page level robots meta tags of noarchive, which it glosses in parentheses as do not use the page for model training, noindex, and none. The same page states that Amazonbot is used to improve Amazon's products and services and may be used to train Amazon AI models, that it respects the Robots Exclusion Protocol by honouring the user-agent and the allow and disallow directives, that it will fetch host level robots.txt files or use a cached copy from the last 30 days, and that it does not support the crawl-delay directive, which is the field this blog found 152 of 1,056 sites setting earlier this month.
The other four vendor pages opened for this post name no page level control at all. OpenAI's bots documentation describes robots.txt and published IP ranges and does not mention a meta tag or an X-Robots-Tag header. Anthropic's crawler support article, carrying 7 April 2026, describes a robots.txt Disallow and the Crawl-delay extension and mentions no meta tag. Perplexity's bots guide describes robots.txt, IP allowlisting and firewall rules and mentions none. Common Crawl's FAQ is the near miss: it says CCBot honours the nofollow attribute as it applies to links embedded on your site, which is a page level signal but about outbound links rather than about this page, and it names no robots meta tag. That is consistent with the narrower finding this blog published separately, that two of nine operator pages name the noindex tag at all.
| Vendor page | Page date it carries | Page level directive it names | Training consequence stated |
|---|---|---|---|
| Amazon, developer.amazon.com/amazonbot | None, footer reads 2010-2026 | noarchive, noindex, none, rel=nofollow | noarchive means do not train |
| Microsoft, blogs.bing.com | 22 September 2023 | NOARCHIVE, NOCACHE | NOARCHIVE excludes from training |
| Google, developers.google.com | Last updated 2026-03-24 UTC | noindex, nosnippet, max-snippet and others | noarchive no longer used at all |
| OpenAI, developers.openai.com | None | None | robots.txt and IP ranges only |
| Anthropic, support.claude.com | 7 April 2026 | None | robots.txt and Crawl-delay only |
| Perplexity, docs.perplexity.ai | None | None | robots.txt, IP allowlist, firewall |
| Common Crawl, commoncrawl.org | None | rel=nofollow on links only | No robots meta tag named |
How many sites send noarchive, and who are they?
Seven of the 1,077 readable home pages carried noarchive, and the composition is the first thing worth noticing: six are news publishers and the seventh is a bank. aftenposten.no, japantimes.co.jp, handelsblatt.com, nzz.ch and nzherald.co.nz carry it in a meta named robots. nikkei.com carries it in a meta named bingbot, which is the only instance in the whole corpus of the directive being aimed at the one vendor whose documentation defines what it does to a chat answer. bbva.com is the bank. Nobody sent nocache, so Microsoft's softer setting, the one that permits a URL, a title and a snippet, has zero takers here.
The exact values matter more than the count, because in six of the seven cases noarchive is sitting inside a list that asks for maximum search presence. handelsblatt.com sends noarchive, max-snippet:-1, max-image-preview:large, max-video-preview:-1, index, follow. bbva.com sends index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1, noarchive. aftenposten.no sends max-image-preview:large, noarchive. These are search optimisation blocks that happen to contain a directive Google has retired, not AI opt out blocks. A reader who assumed the tag signals an editorial decision about model training would be reading intent that the markup does not carry, and this measurement cannot distinguish the two: it records what was delivered, not why.
That is the same discipline this blog applied to the X-Robots-Tag header, where five of 391 home pages sent a noindex and all five turned out to be a captcha rather than a decision about indexing. The directive counts here carry the same caution. Seventeen of the 1,077 sent noindex on what should be a home page, and several of those served almost nothing: nus.edu.sg, statssa.gov.za and zurich.com each returned 212 bytes, rbc.com returned 148, and gymshark.com redirected to us.checkout.gymshark.com before answering. Those are interstitials and shells, and a noindex on them is housekeeping.
| Host | Stratum | Meta name and value delivered | Its robots.txt on Amazonbot |
|---|---|---|---|
| aftenposten.no | news | robots: max-image-preview:large, noarchive | Disallow: / |
| handelsblatt.com | news | robots: noarchive, max-snippet:-1, max-image-preview:large, max-video-preview:-1, index, follow | Disallow: / |
| nzz.ch | news | robots: index, follow, noarchive, noodp | Disallow: / |
| japantimes.co.jp | news | robots: noarchive, in a third separate meta element | No Amazonbot group |
| nzherald.co.nz | news | robots: index, follow, NOODP, noarchive, max-image-preview:large | No Amazonbot group |
| bbva.com | finance | robots: index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1, noarchive | No Amazonbot group |
| nikkei.com | news | bingbot: noarchive | robots.txt not readable |
What the other 434 robots meta tags actually ask for
The census is worth reading in full, because the shape of it contradicts how the meta robots tag is usually discussed. 643 of the 1,077 readable home pages sent no crawler directive of any kind: no meta named robots, no vendor named equivalent, no X-Robots-Tag header. Of the 434 that sent something, 393 sent nothing that withholds anything at all. Their directives were index, follow, all, and the three preview size settings. Only 41 of the 1,077 sent a single directive that restricts or withholds, counting noindex, nofollow, none, noarchive, nosnippet, noimageindex, notranslate, noodp, noydir, noai and noimageai together.
The most common directive in the corpus is max-image-preview, on 234 hosts, and 233 of its 241 occurrences set the value to large. That is the WordPress default rather than an editorial act, which is worth saying plainly because it is the single biggest number in the table and it means almost nothing about intent. max-snippet appears on 133 hosts and takes the value -1, meaning no limit, in 137 of its 142 occurrences. index appears on 311 and follow on 303, both of which are the default behaviour already and change nothing by being stated. The honest summary of the page level surface across this corpus is that where it is used at all, it is overwhelmingly used to ask for more presence rather than less.
Against that, one absence stands out. Not one of the 1,077 pages used any of the fifteen AI crawler tokens in core/src/bots.ts as a meta name. There was no meta named GPTBot, none named ClaudeBot, none named Google-Extended, none named Amazonbot. The only vendor named metas in the entire corpus were googlebot on 23 hosts, bingbot on 7, googlebot-news on 2, msnbot on 1 and adsbot-google on 1. Meanwhile 180 of the 977 hosts that returned both a readable robots.txt and a readable home page named at least one of those fifteen tokens in robots.txt. The AI specific instruction, where it exists at all, lives entirely on the gate and not in the page, which is consistent with what this blog found when it counted that 85 meta elements across five page heads included two that addressed a crawler.
Two sites did reach for an AI specific directive in the page, and they reached for one no vendor page opened for this post mentions. hollywoodreporter.com and variety.com both send noai, noimageai in a meta named robots, alongside separate meta robots elements carrying max-image-preview:large and index, follow. Neither Amazon, Microsoft, Google, OpenAI, Anthropic, Perplexity nor Common Crawl names noai or noimageai on the pages read for this post. That is a statement about seven pages rather than about the whole industry, and it is the bounded version of the claim.
Why a page level opt out needs the crawler to be let in first
This is the finding that makes the other numbers mean something, and it comes from Google's own documentation rather than from an inference. The robots meta tag page states, before any of the directive definitions, that these settings can be read and followed only if crawlers are allowed to access the pages that include these settings. A gate and an instruction compose in one direction only. robots.txt is evaluated against the product token before the request for the page is made, so a Disallow ends the exchange. A meta tag sits in the body of a response that a disallowed crawler never asks for.
Three of the seven sites sending noarchive disallow Amazonbot outright. aftenposten.no, handelsblatt.com and nzz.ch each give Amazonbot a group carrying Disallow: /, which this scanner's own parser resolved using the group selection rules in RFC 9309. For those three, Amazon's crawler never fetches the home page, so it never reads the tag, so the only vendor that documents treating noarchive as a training instruction is the one vendor that has been told to stay outside. The tag is not wrong and it is not harmful. It is unread. Three more, japantimes.co.jp, nzherald.co.nz and bbva.com, name no Amazonbot group at all, which under RFC 9309 group selection means the wildcard group applies if there is one, and those three are the cases where the tag can actually be delivered to Amazon's crawler and acted on. nikkei.com returned a robots.txt that answered 200 with an HTML body, so it was not parseable and the question cannot be answered for that host.
Amazonbot is named as a User-agent line in 107 of the 1,069 readable robots.txt files in the corpus. 47 of those 107 groups carry a Disallow: / and shut it out entirely, 53 disallow some paths but not the root, and 7 allow everything. For comparison, GPTBot is named in 174, with 76 shutting it out at the root, 70 restricting paths and 28 allowing everything, which sits in the same territory as the 82 of 718 robots.txt files this blog measured naming GPTBot in an earlier sweep. The gate is where this corpus does its AI crawler work, by a wide margin, and the page level surface is a rounding error beside it.
There is a practical reading of all this that does not require anybody to change their mind about AI training. If the objective is to keep content out of Amazon's models, the robots.txt Disallow does it on Amazon's own published terms and the noarchive adds nothing on top. If the objective is to stay readable while declining training, the noarchive is the only control either Amazon or Microsoft documents for that, and it only works if the crawler is admitted. Those are different configurations and the corpus contains sites that have assembled parts of both. You can check your own position with the robots.txt tester and against the AI crawlers reference, and read how this scanner classifies what it sees on the methodology page.
Flow: Crawler wants / to Fetches robots.txt; Fetches robots.txt to Matches its own token to a group; Matches its own token to a group (token matched) to Disallow: / applies; Matches its own token to a group (no match) to No rule blocks it; Disallow: / applies to Stops. Page never fetched; No rule blocks it to Fetches the page; Fetches the page to Reads meta robots in the head; Reads meta robots in the head to noarchive can take effect.
Is the tag deliberate, or left over from an older web?
This measurement cannot read intent, and the honest position is to say so and then report the evidence that bears on it. Two of the seven noarchive values sit in the same comma separated list as noodp. nzz.ch sends index, follow, noarchive, noodp. nzherald.co.nz sends index, follow, NOODP, noarchive, max-image-preview:large. noodp and noydir were instructions to Google and Yahoo not to use a third party directory's description in place of the page's own, and Google's current robots meta tag documentation names neither of them anywhere. Fourteen hosts in this corpus still send noodp and eight still send noydir, among them census.gov, medlineplus.gov, otto.de, revolve.com, rakuten.co.jp, kbc.com and progressive.com. A directive block carrying noodp is a block that has not been revisited in a long time, and when noarchive is sitting inside it, the simplest explanation is inheritance rather than a decision about model training.
The rest of the corpus supplies more evidence that these blocks are not closely tended. Seven hosts sent a crawler directive meta with an entirely empty content attribute: seattle.gov, nav.no, falabella.com, flysafair.co.za, allstate.com, becu.org and ocbc.com, the last of which sent two of them. orcid.org shipped the literal string [ROBOTS_PARAMETERS], an unsubstituted template placeholder reaching production in the head of its home page. europa.eu sent noindex follow with no comma between the two, where the specified form is comma separated. bcb.gov.br sent index and follow in a meta named robots while sending noindex in a meta named adsbot-google. mandarinoriental.com sent INDEX,FOLLOW in one meta element and ARCHIVE in another. Fifteen hosts sent more than one meta named robots, which is how japantimes.co.jp ends up declaring index,follow in one element and noarchive in a third.
Seventeen hosts wrote their directives entirely in capitals, including scb.se, cedars-sinai.org, stjude.org, sage.com, bergfreunde.de and heb.com. That one is harmless, since the parsing is case insensitive, and it is listed here only because it is another fingerprint of markup copied forward across years rather than written for the current web. This is the same texture the blog found when it counted defects in robots.txt itself and reported 232 of 1,059 files carrying one. The page level surface is no better maintained than the gate, and it is addressed to an audience that has largely stopped reading it.
-
Empty content attribute7 hosts seattle.gov, nav.no, falabella.com, flysafair.co.za, allstate.com, becu.org, ocbc.com -
Unsubstituted templateorcid.org Shipped the literal string [ROBOTS_PARAMETERS] as its robots directive -
Missing commaeuropa.eu Sent noindex follow where the specified form is comma separated -
Contradiction across metasbcb.gov.br index, follow in robots and noindex in a meta named adsbot-google -
Directive for a dead directory14 and 8 hosts noodp on 14 and noydir on 8, neither named in Google's current documentation -
More than one meta robots15 hosts Directives split across elements, which is how a noarchive ends up isolated
What this measurement does not show
Every figure above is a census of what 1,077 home pages delivered to one client on one day, and that is a narrower thing than it may read as. No crawler was observed. Nothing here is evidence that Amazonbot honoured a noarchive, that Bing excluded a page from a chat answer, or that any model was or was not trained on anything. The vendor rows are each vendor's published claim about itself, read at source on 5 October 2026 and reported as a claim, which is the only standing those statements have here. Where a page carried no date, as Amazon's, OpenAI's and Perplexity's did not, the figures above say so rather than implying freshness.
The scope limits are specific. One page was read per hostname, the home page, so a site that sets noarchive on its articles and not on its front door is invisible to this run, and for a news publisher that is the likelier configuration: this corpus therefore probably undercounts the directive rather than overcounting it. No JavaScript was executed, so a directive injected client side is absent, though a meta robots tag added after load is of doubtful value anyway since the crawlers in question do not all render. The 350 hostnames whose robots.txt was unreadable and the 342 that served no readable home page are absent from every rate, and a host that answers 403 to a plain crawler is exactly the kind of host most likely to hold strong views about crawlers, so the surviving set leans toward sites that admit them. Thirteen hosts disallowed this scanner at the root and were not fetched further, which is a rule this scanner follows and a gap it creates in its own data.
Two further caveats belong to the product rather than the corpus. Lantad does not score any of these directives: the composite grade is built from parity, access, structure and schema, and a clean report from this scanner says nothing at all about whether a page carries noarchive. That is the same gap the earlier post on Microsoft's meta tag opt out named and it has not closed. And the corpus itself is an editorial sampling frame assembled for platform and industry coverage rather than a random draw of the web, so every rate here describes these 1,419 hostnames on one date and nothing wider. The strata are listed on the crawlability study page for anybody who wants to judge the frame before trusting the rates.
Measured on 5 October 2026
- What 1,077 home pages delivered to one client
- Directives in metas and X-Robots-Tag headers
- robots.txt group resolution for 15 AI tokens
- The exact strings each of the 7 sent
Not measured, and not claimed
- Whether any crawler honoured any directive
- Whether any model trained on any page
- Interior pages, which were never fetched
- Why a site sent what it sent
Lantad
Published .
There are two places a site can tell a machine what to do with a page, and they are read at different moments. robots.txt is read before the fetch, which is what makes it a gate. A meta robots tag is read after the fetch, inside the page, which makes it an instruction to a visitor already holding the bytes. The AI crawler conversation has run almost entirely on the first surface: tokens, groups, Disallow lines, and the long argument about whether anybody honours them. The second surface is quieter, and two AI vendors have quietly attached a training consequence to a directive on it.
Common questions
Does noarchive stop AI training?
Only for the two vendors that say it does, and only if their crawler is allowed to fetch the page. Amazon's crawler page states that Amazonbot respects a page level noarchive as do not use the page for model training, and Microsoft's 22 September 2023 post states that content tagged NOARCHIVE will not be used to train its generative AI foundation models and will not appear in Bing Chat answers. OpenAI's, Anthropic's, Perplexity's and Common Crawl's own pages, all read on 5 October 2026, name no robots meta tag at all, so for those four the directive carries no published meaning.
Is noarchive still worth setting if Google ignores it?
Google's robots meta tag documentation, Last updated 2026-03-24 UTC, states that noarchive is no longer used by Google Search because the cached link feature no longer exists, so it costs nothing in Google Search and does nothing there either. Whether it is worth setting depends on the two vendors that do read it, and on whether their crawlers are admitted: Lantad found 3 of the 7 corpus sites sending noarchive had already disallowed Amazonbot outright, which makes the tag unreadable by the crawler that would act on it.
What is the difference between noarchive and nocache?
They differ only for Microsoft, which defines both. Its September 2023 post states that NOARCHIVE content will not be included in Bing Chat answers and will not be used for training, while NOCACHE content may be included with only the URL, title and snippet shown, and only those parts may be used in training. Google's documentation states that neither is used by Google Search. None of the 1,077 readable home pages Lantad measured on 5 October 2026 sent nocache.
Should an AI opt out go in robots.txt or in a meta tag?
robots.txt is read before the page is fetched and a meta tag is read after, so the gate takes precedence and the page level directive only applies to a crawler you have admitted. Google's documentation states that these settings can be read and followed only if crawlers are allowed to access the pages that include them. In Lantad's 5 October 2026 scan the gate is where the work is actually done: 180 of 977 hosts named at least one of fifteen AI crawler tokens in robots.txt, and not one of the 1,077 pages addressed any of those tokens in a meta tag.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.