BlogFindings

Vary: User-Agent: 32 of 1,072 home pages declared it, and 30 of them sent both agents the same bytes

Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt and then for their home page on 6 October 2026, as LantadBot, with no JavaScript executed. 1,072 answered HTTP 200 with HTML. 32 of those declared Vary: User-Agent, 30 of the 32 returned byte identical responses to a crawler user agent string and a desktop Chrome one, and the three pages whose word count actually changed between the two declared no User-Agent at all.

16 min read Lantad

This run counted what sites actually say, rather than what they do. All 1,419 hostnames in this repository's two committed corpus seed files were asked for robots.txt and then for their home page on 6 October 2026, as LantadBot, with redirects followed, a twenty second timeout and no JavaScript executed. 1,072 answered HTTP 200 with HTML. Each of those was requested again with a desktop Chrome user agent string, and every page whose bytes changed was requested twice more under both strings, so that a page which simply differs between two requests is separated from a page that differs by user agent. The declarations and the behaviour turn out to be close to unrelated. The nearest prior run, on conditional requests and what a returning crawler can revalidate against, asked whether a cache could be refreshed at all. This one asks who the stored copy is for.

In short

  • Vary: User-Agent was declared by 32 of the 1,072 corpus home pages Lantad read on 6 October 2026, and 30 of those 32 returned byte identical responses to a crawler user agent string and a desktop Chrome one, so the declaration narrowed a shared cache key without describing any difference in the page.
  • 592 of the 1,072 home pages named Accept-Encoding in Vary and nothing else, 133 sent no Vary header at all, and 91 of those 133 were still publicly cacheable because they sent neither no-store nor private.
  • Of the 1,069 home pages that answered both user agent strings on 6 October 2026, 448 returned different bytes and only 15 did so again when both were requested a second time, so 432 of the 448 raw differences were pages that change between any two requests.
  • The three pages whose word count differed by more than a tenth between the two strings were france24.com, bind.com.mx and nice.org.uk, and each served the LantadBot request a full page while answering the Chrome string with a 403, a bot verification page and a 403, confirmed on a third round the same day.
  • 45 of the 939 Vary values repeated a token, hel.fi naming Accept-Encoding three times, and karolinska.se and psychiatry.org sent Vary: * alongside cache-control: public, a combination RFC 9111 says always fails to match in a shared cache.
What was countedFigureAgainst
Home pages read1,072of 1,419 hostnames asked
Sent a Vary header93987.6 percent of the 1,072
Sent no Vary header13391 of them still publicly cacheable
Named Accept-Encoding and nothing else59255.2 percent of the 1,072
Named User-Agent323.0 percent of the 1,072
Sent Vary: *2both also sent cache-control: public
Repeated a token in the value45of the 939 that sent one
Changed reproducibly with the user agent15of 1,069 that answered both strings
The 1,072 corpus home pages that answered HTTP 200 with HTML, read by Lantad on 6 October 2026. Every row counts what the response headers and bodies contained, not what any AI crawler does with them.

What does Vary: User-Agent do?

The mechanism is small and fully specified, which is why the gap between the specification and the corpus is worth measuring at all. RFC 9111, HTTP Caching, a Standards Track document published in June 2022, puts it in section 4.1, Calculating Cache Keys with the Vary Header Field: when a cache receives a request that could be satisfied by a stored response carrying a Vary header, the cache "MUST NOT use that stored response without revalidation unless all the presented request header fields nominated by that Vary field value match those fields in the original request". MDN's reference for the header, last modified on 21 November 2025, describes the value as the parts of the request message, aside from the method and the URL, that influenced the content of the response.

Two consequences follow, and they pull in opposite directions. Name a header that really does change the response and a shared cache stops mixing the two audiences. Name one that does not and the cache key splits for nothing: two stored copies, two sets of misses, and a lower hit rate bought with no correctness gain. User-Agent is the worst case for that trade because the value space is enormous. Every browser build, every bot and every library version is a distinct key, so a cache told to vary on it is close to a cache told not to store.

Google is the one major search vendor that documents wanting the header. Its page on mobile site and mobile-first indexing best practices, last updated on 10 December 2025, defines dynamic serving as a configuration that uses the same URL regardless of device and says it "relies on user-agent sniffing and the Vary: user-agent HTTP response header to serve a different version of the HTML to different devices". That is the honest frame for this measurement: the header is a declaration about a site's own serving behaviour, and a declaration can be checked against the behaviour it describes. Nothing below is a claim about what any AI crawler does with the header, because no AI vendor documents that, which the last section covers.

How a shared cache decides whether a stored response may be reused, following RFC 9111 section 4.1. A diagram of the specified mechanism, not a measurement of any site.

What 1,072 home pages declared in Vary

939 of the 1,072 pages sent a Vary header and 133 sent none. The distribution inside those 939 is narrow to the point of being a single convention. 885 named Accept-Encoding, and 592 named Accept-Encoding and nothing else, which is 55.2 percent of every page read. That value is almost always not an editorial decision at all: it is what a server or CDN adds when it negotiates gzip or Brotli, and it is correct, because the compressed and uncompressed bodies really are different responses.

After compression the counts fall away fast. 103 pages named Cookie, 85 named Accept, 54 named Origin, and 32 named User-Agent. A long tail of 67 pages named the Next.js App Router set, rsc together with next-router-state-tree, next-router-prefetch and next-router-segment-prefetch, which is the framework distinguishing a React Server Component payload from an HTML document at the same URL. That is a real difference in what comes back and the declaration is accurate, which is worth saying because framework defaults are usually where this blog finds the opposite, as it did when a Next.js payload ran fifteen times the size of its own prose. None of those 67 named User-Agent.

The platform strata in the seed file split cleanly, which suggests the value is set by the hosting layer rather than by the people who write the pages. 32 of the 34 Shopify storefronts sent the identical string "Accept,accept-encoding", a 33rd sent the same two names with a space after the comma, and one sent nothing. 53 of the 54 Wix and Squarespace pages sent Accept-Encoding alone. 19 of the 30 Framer pages sent "Accept-Encoding, Accept". The startup single page applications are the outlier in the other direction: 15 of 44 sent no Vary header at all, the highest share of any stratum, and only 12 sent the compression value alone. Not one page in the government, Wix and Squarespace, Webflow, Framer, Shopify, static documentation, no-code or SaaS marketing strata named User-Agent. The standing method for a single scan is on the methodology page.

  • Accept-Encoding 885 pages compression negotiation, set by the server or CDN
  • Cookie 103 pages a logged in page differs from an anonymous one
  • Accept 85 pages 32 of 34 Shopify storefronts sent this with compression
  • rsc and the Next.js router set 67 pages an RSC payload against an HTML document
  • Origin 54 pages cross origin response headers, not page content
  • User-Agent 32 pages 3.0 percent of the pages read
  • Vary: * 2 pages karolinska.se and psychiatry.org
Request header names appearing in the Vary values of 1,072 corpus home pages, counted once per page, by Lantad on 6 October 2026. A page naming several headers appears in several bars.

The 32 that declared it, and the 30 that did not vary

A declaration is checkable, so the 32 pages naming User-Agent were checked against what they served. Each was requested once as LantadBot and once with a desktop Chrome user agent string, and the two bodies were hashed. 12 of the 32 returned byte identical responses on that first pair. 20 returned something different, which looks at first like the header doing its job.

It is not. Requesting both strings a second time collapses almost all of it. Only two of the 32, flamingoagency.com and mdanderson.org, differed in a way that reproduced: the same crawler body twice and the same browser body twice, reliably distinct from each other. Seventeen of the other 18 differed between the two user agents and also differed between two requests under the same user agent, which means the user agent was not what changed the page, and the eighteenth, sentry.io, could not be judged because its control round failed to return. So 30 of the 32 pages declaring Vary: User-Agent showed no reproducible variation by user agent on 6 October 2026, and the two that did were byte differences rather than content differences: flamingoagency.com served 1,308 words to both and mdanderson.org served 1,539 to both, with the responses differing by 379 and 3,160 bytes of markup around identical prose.

Where the 32 sit is as informative as the count. Finance and ecommerce supplied seven each, travel five, healthcare four, SaaS and news three each, small business WordPress two and education one. Government supplied none of 82. That is the shape of a bot management or personalisation layer rather than a content decision, and it matches what this blog found when it measured 103 of 1,089 home pages refusing a crawler at the CDN. The reader's practical question, which version of a page an engine ends up holding, is the same one behind 29 of 1,087 pages answering a phone and a desktop differently and behind 27 of 719 changing declared language on an Accept-Language header. On this corpus the Vary header is a poor predictor of the answer, and prose parity measured directly is a better one.

StratumNamed User-AgentReproducibly variedVaried in word count
Finance700
Ecommerce700
Travel500
Healthcare410
SaaS300
News300
WordPress, small business210
Education100
Government000
The 32 corpus home pages naming User-Agent in Vary, grouped by seed stratum, with what each group served to a crawler string and a desktop Chrome string. Reproducible means the same distinct bodies on a second round. Measured by Lantad on 6 October 2026.

The three pages whose words changed, and what the Chrome string got

Across the whole corpus, four pages reproducibly returned a different word count to the two user agent strings, and three of those differed by more than a tenth. All three went the opposite way to the direction the word cloaking implies. The page went to the crawler and the refusal went to the browser.

france24.com answered LantadBot with HTTP 200 and 3,060 words of news home page, and answered the Chrome string with HTTP 403 and a 637 byte document headed "Access denied" whose body reads "You don't have permission to access the page you requested" and "The website you are visiting is protected". bind.com.mx answered LantadBot with 1,965 words of Spanish language ERP marketing and answered the Chrome string with HTTP 200 and a 1,963 byte page titled "Bot Verification" carrying a reCAPTCHA form. nice.org.uk answered LantadBot with 632 words and answered the Chrome string with HTTP 403 and a 520 byte nginx error page. Each pair was fetched three times on 6 October 2026 and returned the same split every time. None of the three declared User-Agent in Vary: france24.com sent no Vary header, bind.com.mx sent "Accept-Encoding, Cookie" and nice.org.uk sent "Accept-Encoding".

The honest reading needs the shape of the request stated, because the request is half the result. Both fetches came from one datacentre network location. The browser fetch carried a desktop Chrome user agent string and an Accept header and nothing else: no client hints, no Accept-Language, no cookie, none of the dozen other things a real Chrome sends. A bot management service reading that sees a claimed browser that does not behave like one, which is a stronger signal than the user agent itself, and this blog has measured the same asymmetry from the other side when a headless Chromium was blocked on 15 percent of sites. So the finding is not that these sites hide content from people. It is narrower and still useful: on these three hosts the request that identified itself honestly as a crawler was served the page, the request that claimed to be Chrome was challenged or refused, and no Vary header told any cache on the path that the two would differ. A user agent is a claim rather than an identity, and the terms this crawler works under are on the scanner's bot page.

Sent as LantadBot/1.0

  • HTTP 200, text/html
  • 810,409 bytes
  • 3,060 words of page text
  • title: France 24 - International breaking news
  • no Vary header

Sent as a desktop Chrome string

  • HTTP 403
  • 637 bytes
  • 36 words of page text
  • heading: Access denied
  • no Vary header
france24.com on 6 October 2026, the first of three rounds, requested from the same network location within seconds and differing only in the user agent string. All three rounds produced the same split. Neither response carried a Vary header.

448 pages differed and 432 of them were noise

The control round is the part of this measurement that changed the answer, so it is worth reporting as a finding rather than as method. 1,069 of the 1,072 pages answered both user agent strings. 621 of those returned byte identical bodies. 448 returned different bytes, which is 41.9 percent, and a tool that stopped there would report that four home pages in ten serve a crawler something different from a browser.

One exclusion belongs in the open rather than in a footnote. 56 of the 1,405 hostnames were never reached by this run at all: the network this scan ran from answered for them with its own HTTP 403 carrying an x-block-reason header, so no request left for the site and nothing about those hosts can be read either way. They are mostly large news publishers, among them nytimes.com, arstechnica.com, wired.com, straitstimes.com and indianexpress.com. That leaves 1,349 hostnames actually reached, of which 277 answered with something other than HTTP 200 and HTML or returned no status at all, 157 of those answering 403 and 57 answering 503, and 1,072 were read.

Requesting both strings again leaves 15. 432 of the 448 differences did not reproduce, because the page had changed between two requests under the same user agent as well: a request nonce in a script tag, a CSRF token, a rendered timestamp, a rotating asset hash, a carousel that starts on a different slide. One page could not be controlled because the second round failed to return. 96.4 percent of the raw signal was the page talking to itself, and only the remaining 3.6 percent could be attributed to the user agent at all. Of the 15, eleven returned identical word counts with differing markup, japantimes.co.jp differed by 3 words out of 1,620, shopee.sg returned zero words to both strings because its home page is a client rendered shell, and three were the refusals above.

That ratio is the reason this product will not call a page cloaked on a single pair of fetches, and the reason Lantad withholds a grade rather than guessing when the comparison it needs is not available. An earlier version of the engine compared one crawler body against one browser body and a single changing token was enough to trigger the accusation, which is why the check now requires a minimum body size before it will report suspected cloaking at all. A single URL can be watched the same way with what GPTBot sees, and the general measurement is described under AI visibility. The broader lesson applies to anyone building this comparison: without a repeat round under the same user agent, a user agent diff measures page volatility, not serving behaviour.

From 1,419 seed hostnames to the three pages whose words changed with the user agent. Measured by Lantad on 6 October 2026.

45 repeated a token, two told every cache never to match, and no AI vendor documents any of it

Two smaller defects are worth recording because both are machine detectable and neither is visible in a browser. 45 of the 939 Vary values repeated a token. joincenote.com sent "Accept-Encoding,Accept-Encoding", altitudemarketing.com sent "Accept-Encoding,Accept-Encoding,Cookie", entelect.co.za sent "Accept-Encoding,User-Agent,User-Agent", insee.fr repeated an entire three header CORS group, and hel.fi named Accept-Encoding three times before adding Cookie and X-Consumer-ID. A correct cache collapses duplicates, so the practical cost is close to zero, but the shape is a reliable signature of a value assembled by several layers that do not know about each other, and it travels with the response to every cache on the path.

The second defect is not harmless. karolinska.se and psychiatry.org both sent Vary: *, and both also sent cache-control: public, with max-age=523 and max-age=3600 respectively. RFC 9111 states that a stored response with a Vary value containing a member of "*" always fails to match, so these two responses instruct every shared cache to store them for a set period and then never reuse what it stored. The cacheability directive and the cache key contradict each other. For context on whether a shared cache is in the path at all on this corpus: 459 of the 1,072 responses arrived with an Age header, which a cache sets when it serves a stored response, 569 carried an x-cache or cf-cache-status header, and 362 were served by Cloudflare. Meanwhile 91 of the 133 pages sending no Vary at all were publicly cacheable, sending neither no-store nor private.

What none of this establishes is what an AI crawler does with any of it, and that boundary is the point. OpenAI's crawler documentation names four tokens, OAI-SearchBot, OAI-AdsBot, GPTBot and ChatGPT-User, and says nothing about Vary, HTTP caching, ETag or conditional requests. Anthropic's article on its web crawlers, dated 7 April 2026, names ClaudeBot, Claude-User and Claude-SearchBot and is silent on the same four subjects. The current tokens are listed on the AI crawlers reference. So a site can be measured against the specification, which is what this post did, and cannot be measured against vendor behaviour that is not published. The checkable part is the part worth fixing: declare User-Agent only if the response really depends on it, never pair Vary: * with public cacheability, and make sure the page a crawler gets is the page you meant it to get, which starts with whether it can read your robots.txt at all, a file 92 of 1,056 sites answered with a non-200 when the request carried OpenAI's example GPTBot string.

  • Accept-Encoding Correct and necessary 885 pages. The compressed and uncompressed bodies genuinely are different responses.
  • User-Agent with no variation Splits the key for nothing 30 of the 32 pages naming it served both user agent strings the same bytes.
  • Accept-Encoding,Accept-Encoding Collapses harmlessly 45 values repeated a token. A correct cache dedupes, but the duplication shows several layers writing the header.
  • Vary: * with cache-control: public Stored and never reused karolinska.se and psychiatry.org. RFC 9111 says a stored response whose Vary contains * always fails to match.
Four Vary values read from corpus home pages on 6 October 2026, and what each one asks a shared cache to do. The verdicts follow RFC 9111 section 4.1, not any AI vendor's documented behaviour.

Written by

Lantad

Published .

A shared cache stores one copy of a page and hands it to everyone who asks for the same URL. The header that narrows that promise is Vary, and the only thing it does is add named request headers to the cache key. So Vary: User-Agent is a site telling every cache on the path that the response depended on who asked, and leaving it out is a site telling them the opposite. On a URL that genuinely serves an AI crawler something different from a browser, getting that declaration wrong means a stored copy of one audience's page can be handed to the other.

Common questions

What does Vary: User-Agent actually do?

It adds the User-Agent request header to the cache key, so a shared cache may only reuse a stored response for a later request whose User-Agent matches the original. RFC 9111 section 4.1 is the normative text. It does not change what the origin serves; it describes a dependency the origin already has.

Should I add Vary: User-Agent so AI crawlers get the right page?

Not on its own, and not unless the response genuinely depends on the user agent. No AI vendor documents reading the header: OpenAI's crawler page and Anthropic's crawler article, dated 7 April 2026, both name their tokens and say nothing about Vary or HTTP caching. On the 1,072 corpus home pages Lantad read on 6 October 2026, 30 of the 32 pages declaring it served both user agent strings identical bytes, so the declaration was fragmenting a cache key without describing a difference.

Does a page returning different bytes to a crawler mean it is cloaking?

No, and on this corpus that assumption would be wrong 96 percent of the time. 448 of 1,069 home pages returned different bytes to a crawler string and a Chrome string on 6 October 2026, but only 15 did so again when both were requested a second time. The other 432 also differed between two requests under the same user agent, so what changed was the page, not the audience.

What does Vary: * mean for a shared cache?

RFC 9111 states that a stored response whose Vary value contains a member of * always fails to match, so a shared cache can store the response but can never reuse it. Two corpus pages, karolinska.se and psychiatry.org, sent it alongside cache-control: public with a positive max-age on 6 October 2026, which asks for storage that is then unusable.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.