BlogFindings

Do AI crawlers cache your pages? 391 of 1,101 home pages offered nothing to revalidate against

Lantad requested the home page of all 1,419 hostnames in this repository's two committed corpus seed files twice on 6 October 2026 as LantadBot, then asked a third time with a conditional request carrying the validator the first response had supplied. 1,101 answered HTTP 200 with HTML. 391 of those sent neither an ETag nor a Last-Modified header, so a returning crawler had nothing to ask against, and 514 of the 710 that did offer one answered HTTP 304 Not Modified.

20 min read Lantad

That mechanism only works if the server supplies the token. Google's December 2024 post on HTTP caching put a number on how rarely it does: about 0.026 percent of its total fetches were cacheable ten years before that post, and about 0.017 percent were at the time of writing. That is Google's figure for its own crawl across the whole web, and it describes the outcome rather than the cause. So this run measured the cause, on one page per site, on a fixed corpus. Lantad requested the home page of all 1,419 hostnames in this repository's two committed corpus seed files twice on 6 October 2026, with no JavaScript executed, then sent a third request carrying whichever validator the first response had supplied. 1,101 answered HTTP 200 with an HTML content type. 391 of them, 35.51 percent, supplied nothing to ask against at all.

In short

  • Whether AI crawlers cache your pages is decided by your server rather than by the crawler, and on 6 October 2026 Lantad found that 391 of the 1,101 readable home pages in this repository's committed corpus sent neither an ETag nor a Last-Modified header, so a returning crawler had no cheap way to ask whether anything had changed.
  • Of the 710 corpus home pages that did supply a validator on 6 October 2026, 514 answered a conditional request with HTTP 304 Not Modified, so revalidation worked on 514 of the 1,101 measured pages, which is 46.68 percent of them.
  • 80 of the validators supplied could never match, because 65 ETags and 15 Last-Modified dates changed between two identical requests sent seconds apart, and only one of those 65 ETags still produced a 304.
  • 117 sites sent a validator that did not change and still answered 200 to a conditional request that carried it, among them irs.gov, sec.gov, mit.edu and stanford.edu, each re-confirmed with a second independent request on the same day.
  • Google's December 2024 post on HTTP caching states that about 0.017 percent of its total fetches were cacheable against about 0.026 percent ten years earlier, and of the nine vendors behind the fifteen AI crawler tokens in core/src/bots.ts only Google's own documentation mentions conditional requests at all.
StageHostsWhat happened
Hostnames asked1,419392 in ten platform strata, 1,027 in eight industry strata
Not reached from this network56Refused by this sandbox's own egress policy, not by the site
Never returned a status34Failed in the client, at DNS, or on the timeout
Answered non-200, or 200 without HTML228Mostly 403 and 503 to this crawler
Answered 200 with HTML1,101The denominator for every rate below
Supplied an ETag467240 of the 467 were weak, carrying the W/ prefix
Supplied a Last-Modified474231 sites supplied both headers
Supplied neither391Nothing a returning crawler can ask against
Asked again conditionally710One mechanism per request, never both
Answered HTTP 304 Not Modified51446.68 percent of the 1,101 measured
Supplied a validator that had already changed8065 ETags, 15 dates
Supplied a stable validator and ignored it11789 ETags, 28 dates
Up to three GET requests of https://<host>/ per hostname as LantadBot/1.0 (+https://lantad.co/bot), redirects followed, 20 second timeout, no JavaScript executed, from one network location. Measured by Lantad on 6 October 2026 across the 1,419 hostnames in this repository's two committed corpus seed files.

Do AI crawlers cache your pages, and what actually decides it?

The decision is almost entirely yours, and it is made in two response headers. RFC 9110 defines both halves. A server may send an ETag, an opaque string standing for this exact version of this page, or a Last-Modified date, or both. A client that kept either one may send it back later: the ETag in an If-None-Match request header, the date in an If-Modified-Since one. If the condition says the client already holds the current version, the server answers 304 Not Modified, and section 15.4.5 of that document is explicit about what the answer looks like: a 304 response is terminated by the end of the header section, and it cannot contain content or trailers. No body is the point of the exercise.

What a crawler does with that is published, for exactly one of the vendors this scanner tracks. Google's crawler overview states that Google's crawling infrastructure supports heuristic HTTP caching as defined by the HTTP caching standard, specifically through the ETag and If-None-Match pair and the Last-Modified and If-Modified-Since pair. It states that where a response carries both, Google's crawlers use the ETag value as required by the HTTP standard. It also draws a line that matters for the rest of this post: other HTTP caching directives are not supported. A Cache-Control header is not what Google's crawler reads to decide whether to revalidate. The validator is.

The other eight vendors publish nothing on the subject. The documentation page this scanner records for each of the nine vendors behind the fifteen crawler tokens in core/src/bots.ts was opened on 6 October 2026 and searched for any mention of ETag, If-None-Match, If-Modified-Since, a 304 response or conditional requests. OpenAI's bots documentation mentions none of them: its only occurrences of the word caching refer to prompt caching elsewhere in the API documentation. Anthropic's crawler support article mentions none. Nor does Perplexity's bots guide, nor Amazon's crawler page, nor Common Crawl's CCBot page. Apple's crawler support article, Meta's web crawlers page and ByteDance's webmaster documentation were read the same way and name none of it either, in English or, on the last of those, in Chinese.

That silence is worth stating carefully, because it is an absence of documentation and not an observation of behaviour. Nothing here shows that GPTBot or ClaudeBot fails to send a conditional request. It shows that if one of them does, the site owner has no published statement to rely on, and the request this scan sent as LantadBot is not evidence about any of them. What the measurement below does establish is the half of the exchange that is not in dispute: on 391 of 1,101 home pages, the question could not have been asked by anyone.

The conditional request exchange as RFC 9110 defines it, and the branch 391 of the 1,101 measured home pages took. Mechanism, not a measurement of any crawler's behaviour.

What 1,101 home pages supplied to a returning crawler

The first response from each host was read for the two validators and nothing else. 467 of the 1,101 supplied an ETag, 474 supplied a Last-Modified, and 231 supplied both, which leaves 236 with an ETag alone, 243 with a date alone and 391 with neither. The four groups are exhaustive and they split the corpus almost into quarters, with the empty-handed group the largest single one at 35.51 percent.

That group is not an artefact of one unlucky request. Every one of the 1,101 hosts was asked a second time, and of the 391 that had supplied no validator, 388 supplied none again. Three changed their answer between the two requests, which is a reasonable rate for a corpus of this size and is the kind of detail that decides whether a figure is worth publishing. The sites in that group are not obscure: bls.gov, census.gov, cdc.gov, bundestag.de, europa.eu and gov.scot all served a home page with no ETag and no Last-Modified. For public reference sites whose pages genuinely do sit unchanged for months, that is the most expensive possible configuration, because the content is the most cacheable and the server declines to say so.

Where a validator was present, which one it was matters less than the vendor advice suggests, with one exception. Google's December 2024 post recommends the ETag specifically, on the grounds that it is less prone to errors and mistakes because its value is not structured in the way a date is, and recommends setting both where that is an option. 231 sites took that advice. The structural risk the recommendation names did show up, though barely: of the 474 Last-Modified values delivered, two were not in the format the HTTP specification requires. sonarsource.com sent an ISO 8601 timestamp, 2026-10-02T15:09:08.331Z, and joincenote.com sent a correctly shaped date ending in UTC rather than GMT. Both still parsed here. One more was simply wrong: flipkart.com declared a Last-Modified of Sat, 10 Oct 2026 14:04:34 GMT, four days after the request that received it.

A date that cannot be trusted is a theme this blog keeps meeting from different directions. The same corpus produced 340 of 440 pages whose sitemap lastmod disagreed with the Last-Modified header their own server sent, and separately 38 of 195 pages whose declared dateModified never appeared in the text at all. The difference here is that the Last-Modified header is not an editorial claim about freshness. It is a cache validator, and section 13.1.3 of the HTTP semantics specification says a recipient must ignore If-Modified-Since entirely when the request also carries an If-None-Match. That rule is why each conditional request in this scan carried exactly one mechanism and never both: sending both would have measured the ETag path twice and the date path never, and reported it as having tested two things.

None of this is part of an AI visibility grade, and it is worth saying so before the numbers get any larger. This scanner's composite score is built from parity, access, structure and schema. It does not read a validator, and a page with no ETag and no Last-Modified is penalised nowhere in it.

What the first response carriedHostsShare of 1,101What a returning crawler can send
ETag and Last-Modified23121.0 percentEither; Google's documentation says it uses the ETag
ETag only23621.4 percentIf-None-Match
Last-Modified only24322.1 percentIf-Modified-Since
Neither39135.5 percentNothing, so every revisit is a full download
What the first response carried, across the 1,101 corpus home pages that answered HTTP 200 with HTML. Measured by Lantad on 6 October 2026. The four rows are exhaustive and sum to 1,101.

What happened when the same page was asked again conditionally

710 hosts had supplied something to ask against, and each was asked. Where the first response carried an ETag, the third request sent that exact value in an If-None-Match header and nothing else. Where it carried only a date, the third request sent that date in If-Modified-Since. 514 of the 710 answered HTTP 304 Not Modified. Taken against the measured set rather than the eligible set, revalidation worked on 514 of 1,101 pages, or 46.68 percent, which is the single figure this post exists to establish.

The two mechanisms did not perform equally, and the gap runs the opposite way to the vendor recommendation. 467 ETags produced 314 responses of 304, a rate of 67.2 percent. 243 date-only hosts produced 200 responses of 304, a rate of 82.3 percent. The ETag is the better-specified mechanism and it was the less reliable one in this corpus, for a reason the next section is entirely about. Weak validators, meanwhile, behaved correctly: 240 of the 467 ETags carried the W/ prefix that marks a weak entity tag, and 186 of those 240 answered 304, which is what section 13.1.2 of the specification requires, since a recipient must use the weak comparison function when comparing entity tags for If-None-Match.

One check came back unanimous. Not one of the 514 responses of 304 carried a body. The body was zero bytes on all of them, which is what the specification demands and is the whole economic argument for the mechanism, since the server avoids both generating the page and transferring it. On a corpus where almost every census of delivered markup this blog runs turns up a defect rate in the double digits, a rule obeyed 514 times out of 514 is worth recording.

Partitioning the full 1,101 by outcome gives four groups that account for every page. 391 supplied no validator. 514 supplied one and honoured it. 117 supplied a validator that had not changed between two identical requests and answered 200 to a conditional request carrying it anyway. 79 supplied a validator that had already changed by the time the conditional request went out, so a 200 was the correct answer to a question that could never have succeeded. Those last two groups are 196 pages that look equipped for revalidation and are not, and they are the reason the eligible-set rate of 72.4 percent overstates what is actually happening. This is the same distinction the blog drew when it measured 250 of 647 home pages sending no signal that anything had changed: the presence of a field is not the same as the field working, and the method notes for this scanner treat those as separate results throughout.

  • Answered 304 Not Modified 514 pages 314 by ETag, 200 by date
  • Supplied no validator at all 391 pages 388 of them again on a second request
  • Stable validator, answered 200 anyway 117 pages 89 ETags, 28 dates
  • Validator had already changed 79 pages 64 ETags, 15 dates
Every one of the 1,101 measured home pages in exactly one group, by what it supplied and how it answered a conditional request carrying it. Measured by Lantad on 6 October 2026. The four values sum to 1,101.

The 80 validators that changed between two identical requests

The second plain request to every host existed to test one thing: whether the validator was stable. Two requests seconds apart, nothing changed in between, no JavaScript executed either time. If the ETag or the date differed between them, then the token identifies the response rather than the page, and a conditional request built on it can never succeed. 80 hosts failed that test, 65 on the ETag and 15 on the date. Of the 65 ETag cases, exactly one went on to answer 304, which is best read as a cache node happening to serve the same copy twice rather than as the mechanism working.

33 of the 65 shared one recognisable shape. The ETag began with the string page_cache, followed by a numeric identifier, the string IndexController, and two hexadecimal segments. Across two requests the prefix stayed byte for byte identical and the closing segment moved. allbirds.com supplied page_cache:11044168:IndexController:78b4efa6a0b251a673ed18b4884c6c2d: and then a different closing hash each time. gymshark.com, brooklinen.com, everlane.com, bokksu.com, kammok.com and breda.com did the same. 37 corpus hosts supplied an ETag of that shape in total, 31 of them in the 34 site storefront stratum the seed file assembles from published lists of Shopify stores and 6 in the broader ecommerce stratum. The closing segment changed on 32 of the 37, the prefix was unchanged on 35 of them, and 4 answered 304. This post does not establish which platform or edge layer emits that header, and the honest statement is the narrow one: 32 named sites sent a per-response token in a field whose purpose is to be per-version.

The effect on that stratum is the sharpest contrast in the corpus. 32 of its 34 measured pages supplied an ETag, which is near-universal coverage and better than any industry stratum managed. Only 5 of the 34 answered 304. Validator presence was almost perfect and revalidation almost never worked, which is precisely the failure the second request was added to catch, and a census that counted only header presence would have ranked this group near the top. Anyone running a storefront on a hosted platform can check their own in two curl requests; the per-stack notes for Shopify cover the parts of that stack a site owner can actually change, and this is not one of them.

The remaining 32 churners split between weak tags and bare strings, and several make the mechanism's failure legible at a glance. target.com moved from W/"u1fxvljx2z900h" to W/"zpeaqzhreb900f". yale.edu moved from W/"1791239769-0" to W/"1791243189-0", two unix timestamps 3,420 seconds apart. polytechnique.edu supplied "1791247693" and then "1791247694", a value that had advanced by exactly one between the two requests. coursera.org, frontiersin.org, mpg.de, costco.com and nature.com were in the same group. On the date side, 15 hosts regenerated Last-Modified per response, cshl.edu, nationalgeographic.com, stanfordhealthcare.org, restofworld.org, airbaltic.com and oebb.at among them, and none of those 15 answered 304. A page that reports itself as modified on every request is telling a crawler to download it on every visit, which is the same class of mistake as answering 200 for a URL that does not exist: a technically valid response that destroys the information the field was carrying.

allbirds.com, 6 October 2026, LantadBot

  • GET / HTTP/1.1 200 OK
  • ETag received "page_cache:11044168:IndexController:78b4...6c2d:e4ecdff86e45e4ced2c81e34dbc35a6c"
  • GET / HTTP/1.1 (identical request, seconds later) 200 OK
  • ETag received "page_cache:11044168:IndexController:78b4...6c2d:68f7403601ecb5f5625eeb08e89fb02b"
  • prefix identical, final segment changed the token is per response
  • GET / HTTP/1.1 If-None-Match: the first ETag 200 OK, whole page again
Three requests to one host, in order, as sent and as answered. Measured by Lantad on 6 October 2026. ETag values are reproduced as delivered, abbreviated in the middle where marked.

Which strata answered a conditional request, and which ignored one

Split by how the site is built rather than by how large the organisation is, and the result is the one this blog keeps finding: the hosted site builders are close to perfect and the hand-assembled stacks are not. All 54 of the wix-squarespace pages supplied an ETag and 53 of the 54 answered 304. 40 of 42 webflow pages answered 304, almost all of them on a date rather than an ETag. 27 of 31 framer pages answered. Nobody configured any of that, which is the point: it is a default in a platform, and it is the strongest result in the corpus.

The wordpress-smb stratum is the weakest on coverage, with 23 of its 38 measured pages, 61 percent, supplying no validator at all, and 12 answering 304. The small business sites in it are mostly serving pages that change a few times a year. The ecommerce stratum is weak for both reasons at once: 34 of 65 supplied nothing, 14 of the rest supplied a token that churned, and 9 answered 304. The large public-sector and academic strata sit in between on coverage and do better than their coverage suggests, with government at 42 of 90 and education at 43 of 108.

Then there is the group that is hardest to explain and easiest to verify: 117 sites supplied a validator that did not change between two identical requests, and still returned the whole page when asked conditionally with it. 89 of those were ETags and 28 were dates. Because an unchanged validator makes 304 the correct answer, a 200 here means the conditional header was not evaluated. Four were re-checked the same day with a separate tool, sending the exact ETag the server had just issued, and all four returned 200 again: irs.gov with "1790367357-gzip", sec.gov with "1791237127-gzip", mit.edu with "bbc3-65d0ff941de00" and stanford.edu with W/"4n5pm21ajognob". justice.gov, governo.it, nav.no, anu.edu.au, embl.org, bnf.fr, lmu.de and nus.edu.sg were in the same group, as were edx.org, si.edu, wiley.com, ikea.com, hubspot.com, postman.com and replit.com on the date side.

Cache-Control, by contrast, explains less than it looks like it should. 828 of the 1,101 sent one. 167 sent no-store, 156 sent private, 206 sent no-cache, 636 set a max-age and 303 of those set it to zero. It is tempting to read a no-store as the reason a site returned 200, and for Google's crawler the documentation forecloses that reading: other HTTP caching directives are not supported, so the validator is what gets read. 50 sites sent a no-store alongside a perfectly good validator, which is a header block asking for two different things. For anyone who wants to look at their own, the robots.txt tester answers the access question that comes before any of this, and a site that is refusing a crawler it believes it allows has a larger problem than an absent ETag. The crawl-efficiency argument that ends in a Disallow line usually skips this header pair entirely, and it is the cheapest thing on the list.

StratumPages measuredSupplied no validatorAnswered 304
wix-squarespace54053
webflow42240
framer31327
static-docs31724
saas-marketing38925
bubble-nocode421423
spa-startups441623
media-local31818
saas1224657
education1084443
government904042
healthcare954539
finance943433
travel793526
news632915
wordpress-smb382312
ecommerce65349
shopify-dtc3425
Every stratum in the corpus, by what its measured home pages supplied and how many answered a conditional request with 304. Measured by Lantad on 6 October 2026. The pages column sums to 1,101.

What this measurement does not show

No crawler was observed. Every request in this scan was sent by this scanner as LantadBot from one network location, and nothing here is evidence that GPTBot, ClaudeBot, PerplexityBot or any other bot sends a conditional request, honours a 304, or revisits at any particular rate. The vendor documentation finding is narrower still: it records that eight of the nine vendor pages this scanner cites do not mention the mechanism, which is a fact about their documentation and not about their crawlers. Google's figure of about 0.017 percent of fetches being cacheable is Google's own published measurement of its own crawl, quoted from its December 2024 post, and it is not comparable to any rate computed here, because this scan measures server configuration on 1,101 home pages and that figure measures a share of all Google fetches across the web.

One page was read per hostname and it was the home page, which is the page most likely to be dynamic and least likely to be a good caching candidate on any site. A corpus of article pages would very probably produce different and better numbers, and the figures here should not be read as a site-wide property. The three requests to each host were sent within a few seconds of each other, which is a narrow window: a validator that is stable across two requests seconds apart may still churn across a day, so 80 is a floor on the churn rate rather than an estimate of it, and a site that deploys between two requests would appear in that group wrongly.

The partition into stable-but-ignored and already-changed rests on comparing two plain requests, and that is what makes it reportable. Without it, a 200 answer to If-Modified-Since is ambiguous, because the page genuinely may have changed. cshl.edu is the case that forced the distinction: it answered 200 during the scan and 304 to a later check, because it regenerates Last-Modified on every response, and calling it a server that ignores conditional requests would have been wrong. 56 hostnames were never reached at all, refused by this sandbox's own network egress policy rather than by the sites, and they are absent from every figure above. 53 of those 56 sit in the news stratum and most are large international publishers, so the news result here rests on 63 pages rather than the 128 hostnames the seed file provides.

Finally, the corpus is an editorial sampling frame assembled for platform and industry coverage rather than a random draw of the web, so every number describes these 1,419 hostnames on one date and nothing wider. And this is not something the scanner grades: there is no caching signal in the composite score, so a clean report from what GPTBot sees says nothing at all about whether the page can be revalidated. That gap is real and it is stated here rather than quietly closed, which is the same reason this blog publishes the measurements that are inconvenient for the product.

Measured on 6 October 2026

  • What 1,101 home pages supplied in two response headers
  • Whether the validator was stable across two identical requests
  • How each server answered one conditional request carrying it
  • What nine vendor documentation pages say about the mechanism
  • Cache-Control directives as delivered

Not measured, and not claimed

  • Whether any AI crawler sends a conditional request
  • Whether any crawler honours a 304 it receives
  • How often any crawler revisits any page
  • Anything about pages other than the home page
  • Churn over a day rather than over a few minutes
The boundary of this measurement, stated before anybody quotes a figure from it.

Written by

Lantad

Published .

Almost every measurement on this blog asks what a crawler gets on its first visit. This one asks what it gets on its second. An AI crawler that has already read your home page does not have to download it again to find out whether it changed: HTTP has carried a cheap way to ask that question since long before any of these bots existed. The crawler keeps a small token from the previous response, sends it back on the next request, and the server either says nothing has changed, in four hundred bytes of headers and no body, or sends the page. One round trip, almost no bandwidth, and the crawler learns the page is current.

Common questions

Do AI crawlers cache your pages?

Only if your server gives them something to cache against, and most of the vendors do not say whether they try. A crawler can confirm a page is unchanged by sending back an ETag or a Last-Modified date the server supplied earlier, and getting HTTP 304 Not Modified in reply. Lantad measured the server half of that exchange on 6 October 2026 across 1,101 readable corpus home pages: 391 supplied neither header, so no crawler could have asked. Of the nine vendors behind the fifteen crawler tokens in core/src/bots.ts, only Google's documentation states that its crawling infrastructure supports conditional requests at all.

Should I send an ETag or a Last-Modified header?

Google's December 2024 post on HTTP caching recommends the ETag, because its value is not structured in the way a date is and so is less prone to errors, and recommends setting both where you can. Its crawler overview adds that where a response carries both, Google's crawlers use the ETag. In Lantad's 6 October 2026 scan the date was the more reliable mechanism in practice, with 200 of 243 date-only hosts answering 304 against 314 of 467 ETag hosts, because a large group of sites issue an ETag that changes on every response.

Why would a server send an ETag and still return the whole page?

Because the conditional header was not evaluated. If the ETag has not changed, HTTP 304 is the correct answer, so a 200 means the If-None-Match was ignored somewhere in the stack. Lantad found 117 such sites on 6 October 2026, 89 on an ETag and 28 on a date, and re-confirmed four of them the same day with a separate tool using the exact ETag just issued: irs.gov, sec.gov, mit.edu and stanford.edu each returned 200 again. The other common cause is different: the validator did change, because the server regenerates it per response, which happened on 80 hosts.

Does a no-store Cache-Control header stop a crawler revalidating?

Not for Google, whose crawler overview states plainly that other HTTP caching directives are not supported, so the ETag and Last-Modified validators are what its crawlers read. 828 of the 1,101 home pages Lantad measured on 6 October 2026 sent a Cache-Control header, 167 of them no-store, and 50 sites sent a no-store alongside a working validator. No other vendor in this scanner's registry publishes a position on the question either way, so for the rest it is undocumented rather than settled.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.