BlogFindings

noarchive: 7 of 1,077 home pages sent the tag two AI vendors read as a training opt out

Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt on 5 October 2026 and read every permitted home page with no JavaScript executed. 1,077 answered HTTP 200 with HTML. 7 of them sent noarchive, the page level directive Amazon's own crawler page reads as do not train and Microsoft's reads as exclusion from Bing Chat answers, and Google's current documentation says it no longer uses at all. None sent nocache.

19 min read Lantad

The directive is noarchive. It was built for a feature that no longer exists, the cached copy of a page a search engine would show beside the result, and Google now says so in plain terms. Amazon and Microsoft both read it as something else. So this run counted it. Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt on 5 October 2026, evaluated each file for this scanner's own token, and read the home page of every host that allowed it with no JavaScript executed and no stylesheet fetched. 1,077 answered HTTP 200 with an HTML content type. Seven of those 1,077 sent noarchive. The number is small, and the interesting part is not the number: it is that three of the seven have already shut the door on the one crawler that would act on it.

In short

  • noarchive is a meta robots directive from the era of the search engine cached page, and on 5 October 2026 Lantad found it on 7 of 1,077 readable home pages in this repository's committed corpus. None of the 1,077 sent nocache.
  • Amazon's crawler page states that Amazonbot respects a page level noarchive as do not use the page for model training. Microsoft's 22 September 2023 post states that content tagged NOARCHIVE will not be included in Bing Chat answers and will not be used to train its generative AI foundation models.
  • Google's robots meta tag documentation, Last updated 2026-03-24 UTC, states that noarchive is no longer used by Google Search because the cached link feature no longer exists, and that nocache is not used by Google Search either.
  • Three of the seven sites sending noarchive already disallow Amazonbot outright in robots.txt, so Amazon's crawler never fetches the page that carries the tag. Google's own documentation states these settings can be read and followed only if crawlers are allowed to access the pages that include them.
  • Not one of the 1,077 pages addressed any of the fifteen AI crawler tokens in core/src/bots.ts by name in a meta tag, while 180 of the 977 hosts that returned both a readable robots.txt and a readable home page named at least one of them in robots.txt.
StageHostsWhat happened
Hostnames asked1,419392 in ten platform strata, 1,027 in eight industry strata
Returned a parseable robots.txt1,069HTTP 200 with a body that was not HTML; 10 of them empty
Disallowed this crawler at the root13Not fetched further
Answered HTTP 403 to the home page215Refused this crawler
Answered HTTP 503 to the home page56Served no page to this client
Never returned a status34Failed in the client, at DNS, or on the timeout
Answered some other status239 of 429, 5 of 202, 4 of 404, 2 of 406 and one each of 401, 451, 498 and 500
Answered 200 with HTML1,077The denominator for every rate below
Sent any crawler directive434432 in a meta named robots, 25 in a vendor named meta, 12 in an X-Robots-Tag header
Sent noarchive76 news sites and one bank
Sent nocache0Nobody
One GET of https://<host>/robots.txt and then of https://<host>/ per hostname as LantadBot/1.0 (+https://lantad.co/bot), redirects followed, 20 second timeout, no JavaScript executed, from one network location. Measured by Lantad on 5 October 2026 across the 1,419 hostnames in this repository's two committed corpus seed files.

What does noarchive do for an AI crawler now?

Three vendors publish an answer and the three answers disagree. Google's robots meta tag documentation, carrying Last updated 2026-03-24 UTC when it was opened on 5 October 2026, is the shortest: the noarchive rule is no longer used by Google Search to control whether a cached link is shown in search results, as the cached link feature no longer exists. The same page says the nocache rule is not used by Google Search. Both directives are, from Google's side, dead letters on a page that still carefully specifies noindex, nosnippet, max-snippet, max-image-preview and the rest.

Microsoft's reading is the opposite and it is older. Microsoft's September 2023 post on controlling content in Bing Chat, dated 22 September 2023, states that content tagged NOARCHIVE will not be included in Bing Chat answers and will not be linked to in the answers, and that Microsoft will not use that content for training its generative AI foundation models. NOCACHE is the softer setting in the same post: content carrying it may be included in Bing Chat answers with only URL, snippet and title displayed, and only URLs, titles and snippets may be used in training. This blog has written about that pair before, in the post on how Microsoft's AI opt out is a meta tag rather than a robots.txt token, and said at the time that most sites would carry neither without measuring whether that was true. This run measures it.

Amazon's own crawler page is the third reading and the most direct about training. It states that Amazonbot respects the link level rel=nofollow directive and page level robots meta tags of noarchive, which it glosses in parentheses as do not use the page for model training, noindex, and none. The same page states that Amazonbot is used to improve Amazon's products and services and may be used to train Amazon AI models, that it respects the Robots Exclusion Protocol by honouring the user-agent and the allow and disallow directives, that it will fetch host level robots.txt files or use a cached copy from the last 30 days, and that it does not support the crawl-delay directive, which is the field this blog found 152 of 1,056 sites setting earlier this month.

The other four vendor pages opened for this post name no page level control at all. OpenAI's bots documentation describes robots.txt and published IP ranges and does not mention a meta tag or an X-Robots-Tag header. Anthropic's crawler support article, carrying 7 April 2026, describes a robots.txt Disallow and the Crawl-delay extension and mentions no meta tag. Perplexity's bots guide describes robots.txt, IP allowlisting and firewall rules and mentions none. Common Crawl's FAQ is the near miss: it says CCBot honours the nofollow attribute as it applies to links embedded on your site, which is a page level signal but about outbound links rather than about this page, and it names no robots meta tag. That is consistent with the narrower finding this blog published separately, that two of nine operator pages name the noindex tag at all.

Vendor pagePage date it carriesPage level directive it namesTraining consequence stated
Amazon, developer.amazon.com/amazonbotNone, footer reads 2010-2026noarchive, noindex, none, rel=nofollownoarchive means do not train
Microsoft, blogs.bing.com22 September 2023NOARCHIVE, NOCACHENOARCHIVE excludes from training
Google, developers.google.comLast updated 2026-03-24 UTCnoindex, nosnippet, max-snippet and othersnoarchive no longer used at all
OpenAI, developers.openai.comNoneNonerobots.txt and IP ranges only
Anthropic, support.claude.com7 April 2026Nonerobots.txt and Crawl-delay only
Perplexity, docs.perplexity.aiNoneNonerobots.txt, IP allowlist, firewall
Common Crawl, commoncrawl.orgNonerel=nofollow on links onlyNo robots meta tag named
What each vendor's own page says about page level directives, read at source on 5 October 2026. Reported from the vendors' published claims, not measured by Lantad: no crawler was observed honouring or ignoring any of these.

How many sites send noarchive, and who are they?

Seven of the 1,077 readable home pages carried noarchive, and the composition is the first thing worth noticing: six are news publishers and the seventh is a bank. aftenposten.no, japantimes.co.jp, handelsblatt.com, nzz.ch and nzherald.co.nz carry it in a meta named robots. nikkei.com carries it in a meta named bingbot, which is the only instance in the whole corpus of the directive being aimed at the one vendor whose documentation defines what it does to a chat answer. bbva.com is the bank. Nobody sent nocache, so Microsoft's softer setting, the one that permits a URL, a title and a snippet, has zero takers here.

The exact values matter more than the count, because in six of the seven cases noarchive is sitting inside a list that asks for maximum search presence. handelsblatt.com sends noarchive, max-snippet:-1, max-image-preview:large, max-video-preview:-1, index, follow. bbva.com sends index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1, noarchive. aftenposten.no sends max-image-preview:large, noarchive. These are search optimisation blocks that happen to contain a directive Google has retired, not AI opt out blocks. A reader who assumed the tag signals an editorial decision about model training would be reading intent that the markup does not carry, and this measurement cannot distinguish the two: it records what was delivered, not why.

That is the same discipline this blog applied to the X-Robots-Tag header, where five of 391 home pages sent a noindex and all five turned out to be a captcha rather than a decision about indexing. The directive counts here carry the same caution. Seventeen of the 1,077 sent noindex on what should be a home page, and several of those served almost nothing: nus.edu.sg, statssa.gov.za and zurich.com each returned 212 bytes, rbc.com returned 148, and gymshark.com redirected to us.checkout.gymshark.com before answering. Those are interstitials and shells, and a noindex on them is housekeeping.

HostStratumMeta name and value deliveredIts robots.txt on Amazonbot
aftenposten.nonewsrobots: max-image-preview:large, noarchiveDisallow: /
handelsblatt.comnewsrobots: noarchive, max-snippet:-1, max-image-preview:large, max-video-preview:-1, index, followDisallow: /
nzz.chnewsrobots: index, follow, noarchive, noodpDisallow: /
japantimes.co.jpnewsrobots: noarchive, in a third separate meta elementNo Amazonbot group
nzherald.co.nznewsrobots: index, follow, NOODP, noarchive, max-image-preview:largeNo Amazonbot group
bbva.comfinancerobots: index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1, noarchiveNo Amazonbot group
nikkei.comnewsbingbot: noarchiverobots.txt not readable
Every home page in the corpus carrying noarchive, with the exact meta value delivered and what the same site's robots.txt says about Amazonbot, the one crawler whose documentation reads the directive as a training opt out. Measured by Lantad on 5 October 2026.

What the other 434 robots meta tags actually ask for

The census is worth reading in full, because the shape of it contradicts how the meta robots tag is usually discussed. 643 of the 1,077 readable home pages sent no crawler directive of any kind: no meta named robots, no vendor named equivalent, no X-Robots-Tag header. Of the 434 that sent something, 393 sent nothing that withholds anything at all. Their directives were index, follow, all, and the three preview size settings. Only 41 of the 1,077 sent a single directive that restricts or withholds, counting noindex, nofollow, none, noarchive, nosnippet, noimageindex, notranslate, noodp, noydir, noai and noimageai together.

The most common directive in the corpus is max-image-preview, on 234 hosts, and 233 of its 241 occurrences set the value to large. That is the WordPress default rather than an editorial act, which is worth saying plainly because it is the single biggest number in the table and it means almost nothing about intent. max-snippet appears on 133 hosts and takes the value -1, meaning no limit, in 137 of its 142 occurrences. index appears on 311 and follow on 303, both of which are the default behaviour already and change nothing by being stated. The honest summary of the page level surface across this corpus is that where it is used at all, it is overwhelmingly used to ask for more presence rather than less.

Against that, one absence stands out. Not one of the 1,077 pages used any of the fifteen AI crawler tokens in core/src/bots.ts as a meta name. There was no meta named GPTBot, none named ClaudeBot, none named Google-Extended, none named Amazonbot. The only vendor named metas in the entire corpus were googlebot on 23 hosts, bingbot on 7, googlebot-news on 2, msnbot on 1 and adsbot-google on 1. Meanwhile 180 of the 977 hosts that returned both a readable robots.txt and a readable home page named at least one of those fifteen tokens in robots.txt. The AI specific instruction, where it exists at all, lives entirely on the gate and not in the page, which is consistent with what this blog found when it counted that 85 meta elements across five page heads included two that addressed a crawler.

Two sites did reach for an AI specific directive in the page, and they reached for one no vendor page opened for this post mentions. hollywoodreporter.com and variety.com both send noai, noimageai in a meta named robots, alongside separate meta robots elements carrying max-image-preview:large and index, follow. Neither Amazon, Microsoft, Google, OpenAI, Anthropic, Perplexity nor Common Crawl names noai or noimageai on the pages read for this post. That is a statement about seven pages rather than about the whole industry, and it is the bounded version of the claim.

  • index 311 hosts Already the default
  • follow 303 hosts Already the default
  • max-image-preview 234 hosts 233 of 241 occurrences set large
  • max-snippet 133 hosts 137 of 142 occurrences set -1
  • max-video-preview 133 hosts 133 of 139 occurrences set -1
  • noindex 17 hosts Several served under 1,200 bytes
  • noodp 14 hosts For a directory Google's page no longer names
  • all 13 hosts The explicit no restrictions value
  • nofollow 10 hosts Honoured by CCBot and Amazonbot per their own pages
  • noydir 8 hosts For a directory Google's page no longer names
  • noarchive 7 hosts The directive this post is about
  • archive 3 hosts The affirmative form, on cdc.gov, takealot.com, mandarinoriental.com
  • noai and noimageai 2 hosts hollywoodreporter.com and variety.com
  • noimageindex 2 hosts gov.br and iitb.ac.in
  • notranslate 2 hosts iitb.ac.in and siigo.com
  • nosnippet 1 hosts iitb.ac.in
  • nocache 0 Microsoft's softer setting has no takers here
Every crawler directive found across the 1,077 readable home pages, counted by host rather than by occurrence, from a meta named robots, a vendor named meta, or an X-Robots-Tag header. Measured by Lantad on 5 October 2026.

Why a page level opt out needs the crawler to be let in first

This is the finding that makes the other numbers mean something, and it comes from Google's own documentation rather than from an inference. The robots meta tag page states, before any of the directive definitions, that these settings can be read and followed only if crawlers are allowed to access the pages that include these settings. A gate and an instruction compose in one direction only. robots.txt is evaluated against the product token before the request for the page is made, so a Disallow ends the exchange. A meta tag sits in the body of a response that a disallowed crawler never asks for.

Three of the seven sites sending noarchive disallow Amazonbot outright. aftenposten.no, handelsblatt.com and nzz.ch each give Amazonbot a group carrying Disallow: /, which this scanner's own parser resolved using the group selection rules in RFC 9309. For those three, Amazon's crawler never fetches the home page, so it never reads the tag, so the only vendor that documents treating noarchive as a training instruction is the one vendor that has been told to stay outside. The tag is not wrong and it is not harmful. It is unread. Three more, japantimes.co.jp, nzherald.co.nz and bbva.com, name no Amazonbot group at all, which under RFC 9309 group selection means the wildcard group applies if there is one, and those three are the cases where the tag can actually be delivered to Amazon's crawler and acted on. nikkei.com returned a robots.txt that answered 200 with an HTML body, so it was not parseable and the question cannot be answered for that host.

Amazonbot is named as a User-agent line in 107 of the 1,069 readable robots.txt files in the corpus. 47 of those 107 groups carry a Disallow: / and shut it out entirely, 53 disallow some paths but not the root, and 7 allow everything. For comparison, GPTBot is named in 174, with 76 shutting it out at the root, 70 restricting paths and 28 allowing everything, which sits in the same territory as the 82 of 718 robots.txt files this blog measured naming GPTBot in an earlier sweep. The gate is where this corpus does its AI crawler work, by a wide margin, and the page level surface is a rounding error beside it.

There is a practical reading of all this that does not require anybody to change their mind about AI training. If the objective is to keep content out of Amazon's models, the robots.txt Disallow does it on Amazon's own published terms and the noarchive adds nothing on top. If the objective is to stay readable while declining training, the noarchive is the only control either Amazon or Microsoft documents for that, and it only works if the crawler is admitted. Those are different configurations and the corpus contains sites that have assembled parts of both. You can check your own position with the robots.txt tester and against the AI crawlers reference, and read how this scanner classifies what it sees on the methodology page.

The order in which the two surfaces are read, which is why a Disallow makes a page level directive unreachable. Drawn from the mechanism Google's robots meta tag documentation states and the group selection rules in RFC 9309; this is an explanation of the protocol rather than a measurement of any crawler.

Is the tag deliberate, or left over from an older web?

This measurement cannot read intent, and the honest position is to say so and then report the evidence that bears on it. Two of the seven noarchive values sit in the same comma separated list as noodp. nzz.ch sends index, follow, noarchive, noodp. nzherald.co.nz sends index, follow, NOODP, noarchive, max-image-preview:large. noodp and noydir were instructions to Google and Yahoo not to use a third party directory's description in place of the page's own, and Google's current robots meta tag documentation names neither of them anywhere. Fourteen hosts in this corpus still send noodp and eight still send noydir, among them census.gov, medlineplus.gov, otto.de, revolve.com, rakuten.co.jp, kbc.com and progressive.com. A directive block carrying noodp is a block that has not been revisited in a long time, and when noarchive is sitting inside it, the simplest explanation is inheritance rather than a decision about model training.

The rest of the corpus supplies more evidence that these blocks are not closely tended. Seven hosts sent a crawler directive meta with an entirely empty content attribute: seattle.gov, nav.no, falabella.com, flysafair.co.za, allstate.com, becu.org and ocbc.com, the last of which sent two of them. orcid.org shipped the literal string [ROBOTS_PARAMETERS], an unsubstituted template placeholder reaching production in the head of its home page. europa.eu sent noindex follow with no comma between the two, where the specified form is comma separated. bcb.gov.br sent index and follow in a meta named robots while sending noindex in a meta named adsbot-google. mandarinoriental.com sent INDEX,FOLLOW in one meta element and ARCHIVE in another. Fifteen hosts sent more than one meta named robots, which is how japantimes.co.jp ends up declaring index,follow in one element and noarchive in a third.

Seventeen hosts wrote their directives entirely in capitals, including scb.se, cedars-sinai.org, stjude.org, sage.com, bergfreunde.de and heb.com. That one is harmless, since the parsing is case insensitive, and it is listed here only because it is another fingerprint of markup copied forward across years rather than written for the current web. This is the same texture the blog found when it counted defects in robots.txt itself and reported 232 of 1,059 files carrying one. The page level surface is no better maintained than the gate, and it is addressed to an audience that has largely stopped reading it.

  • Empty content attribute 7 hosts seattle.gov, nav.no, falabella.com, flysafair.co.za, allstate.com, becu.org, ocbc.com
  • Unsubstituted template orcid.org Shipped the literal string [ROBOTS_PARAMETERS] as its robots directive
  • Missing comma europa.eu Sent noindex follow where the specified form is comma separated
  • Contradiction across metas bcb.gov.br index, follow in robots and noindex in a meta named adsbot-google
  • Directive for a dead directory 14 and 8 hosts noodp on 14 and noydir on 8, neither named in Google's current documentation
  • More than one meta robots 15 hosts Directives split across elements, which is how a noarchive ends up isolated
Defects and fossils found in crawler directive metas across the 1,077 readable home pages, each named with the host it was found on. Measured by Lantad on 5 October 2026.

What this measurement does not show

Every figure above is a census of what 1,077 home pages delivered to one client on one day, and that is a narrower thing than it may read as. No crawler was observed. Nothing here is evidence that Amazonbot honoured a noarchive, that Bing excluded a page from a chat answer, or that any model was or was not trained on anything. The vendor rows are each vendor's published claim about itself, read at source on 5 October 2026 and reported as a claim, which is the only standing those statements have here. Where a page carried no date, as Amazon's, OpenAI's and Perplexity's did not, the figures above say so rather than implying freshness.

The scope limits are specific. One page was read per hostname, the home page, so a site that sets noarchive on its articles and not on its front door is invisible to this run, and for a news publisher that is the likelier configuration: this corpus therefore probably undercounts the directive rather than overcounting it. No JavaScript was executed, so a directive injected client side is absent, though a meta robots tag added after load is of doubtful value anyway since the crawlers in question do not all render. The 350 hostnames whose robots.txt was unreadable and the 342 that served no readable home page are absent from every rate, and a host that answers 403 to a plain crawler is exactly the kind of host most likely to hold strong views about crawlers, so the surviving set leans toward sites that admit them. Thirteen hosts disallowed this scanner at the root and were not fetched further, which is a rule this scanner follows and a gap it creates in its own data.

Two further caveats belong to the product rather than the corpus. Lantad does not score any of these directives: the composite grade is built from parity, access, structure and schema, and a clean report from this scanner says nothing at all about whether a page carries noarchive. That is the same gap the earlier post on Microsoft's meta tag opt out named and it has not closed. And the corpus itself is an editorial sampling frame assembled for platform and industry coverage rather than a random draw of the web, so every rate here describes these 1,419 hostnames on one date and nothing wider. The strata are listed on the crawlability study page for anybody who wants to judge the frame before trusting the rates.

Measured on 5 October 2026

  • What 1,077 home pages delivered to one client
  • Directives in metas and X-Robots-Tag headers
  • robots.txt group resolution for 15 AI tokens
  • The exact strings each of the 7 sent

Not measured, and not claimed

  • Whether any crawler honoured any directive
  • Whether any model trained on any page
  • Interior pages, which were never fetched
  • Why a site sent what it sent
What this run measured against what it did not, stated so the figures are not read as more than they are.

Written by

Lantad

Published .

There are two places a site can tell a machine what to do with a page, and they are read at different moments. robots.txt is read before the fetch, which is what makes it a gate. A meta robots tag is read after the fetch, inside the page, which makes it an instruction to a visitor already holding the bytes. The AI crawler conversation has run almost entirely on the first surface: tokens, groups, Disallow lines, and the long argument about whether anybody honours them. The second surface is quieter, and two AI vendors have quietly attached a training consequence to a directive on it.

Common questions

Does noarchive stop AI training?

Only for the two vendors that say it does, and only if their crawler is allowed to fetch the page. Amazon's crawler page states that Amazonbot respects a page level noarchive as do not use the page for model training, and Microsoft's 22 September 2023 post states that content tagged NOARCHIVE will not be used to train its generative AI foundation models and will not appear in Bing Chat answers. OpenAI's, Anthropic's, Perplexity's and Common Crawl's own pages, all read on 5 October 2026, name no robots meta tag at all, so for those four the directive carries no published meaning.

Is noarchive still worth setting if Google ignores it?

Google's robots meta tag documentation, Last updated 2026-03-24 UTC, states that noarchive is no longer used by Google Search because the cached link feature no longer exists, so it costs nothing in Google Search and does nothing there either. Whether it is worth setting depends on the two vendors that do read it, and on whether their crawlers are admitted: Lantad found 3 of the 7 corpus sites sending noarchive had already disallowed Amazonbot outright, which makes the tag unreadable by the crawler that would act on it.

What is the difference between noarchive and nocache?

They differ only for Microsoft, which defines both. Its September 2023 post states that NOARCHIVE content will not be included in Bing Chat answers and will not be used for training, while NOCACHE content may be included with only the URL, title and snippet shown, and only those parts may be used in training. Google's documentation states that neither is used by Google Search. None of the 1,077 readable home pages Lantad measured on 5 October 2026 sent nocache.

Should an AI opt out go in robots.txt or in a meta tag?

robots.txt is read before the page is fetched and a meta tag is read after, so the gate takes precedence and the page level directive only applies to a crawler you have admitted. Google's documentation states that these settings can be read and followed only if crawlers are allowed to access the pages that include them. In Lantad's 5 October 2026 scan the gate is where the work is actually done: 180 of 977 hosts named at least one of fifteen AI crawler tokens in robots.txt, and not one of the 1,077 pages addressed any of those tokens in a meta tag.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.