BlogFindings

Do AI crawlers read hidden text? 121,312 of 1,269,054 words sat behind a marker on 1,115 home pages

Lantad requested the robots.txt and then the home page of all 1,419 hostnames in this repository's committed corpus on 30 September 2026 and read the delivered bytes with no JavaScript executed and no stylesheet fetched. 1,115 pages were both allowed and readable, and they carried 1,269,054 words of body text. 121,312 of those words, 9.56 percent, sat inside an element marked hidden from somebody: from a sighted visitor, or from assistive technology, or from both. On 77 of the 820 pages that declared an h1, every h1 was inside one of those elements.

24 min read Lantad

This run counted the difference. Lantad requested the robots.txt of all 1,419 hostnames in this repository's two committed corpus seed files on 30 September 2026, then the home page of each, sent as its own declared crawler user agent with redirects followed and a twenty second timeout, from one network location, with no JavaScript executed and no stylesheet fetched. 1,125 answered HTTP 200 with an HTML content type. 12 of the 1,419 robots.txt files disallowed this scanner at the site root, and 10 of those 12 were among the 1,125, so they were dropped rather than read. That leaves 1,115 pages, carrying 1,269,054 words of body text, and the question is how many of those words an AI crawler would find that a person would not.

In short

  • Do AI crawlers read hidden text is answerable in one direction only: Lantad read 1,115 home pages on 30 September 2026 and found 121,312 of their 1,269,054 body words inside an element marked hidden from somebody, but no crawler operator publishes what its extractor does with any of the four markers counted here.
  • 623 of the 1,115 pages carried at least one such word on 30 September 2026, and the marker was not evenly used: aria-hidden held 65,977 words across 575 pages, an inline display none style held 20,020 across 244, the HTML hidden attribute held 19,724 across 232, and a screen reader only class name held 12,476 across 403.
  • 77 of the 820 home pages that declared an h1 put every one of them inside a marked element, harvard.edu, mit.edu, time.com and usps.com among them, and 66 of those 77 used a screen reader only class name rather than a marker that is exact in the delivered bytes.
  • 12,393 of the 19,802 marked regions sat outside any nav, header, footer or aside element and outside any navigation, banner, contentinfo, menubar or menu role, carrying 74,897 words, so the extra text is not all duplicated menus even though most of the largest single regions are.
  • The Framer stratum carried the highest share at 8,443 of 39,473 words on 31 pages, and the static documentation stratum the lowest at 683 of 29,138 on 31 pages, measured by Lantad on 30 September 2026.
StageCountWhat happened
Hostnames requested1,419The committed corpus, an editorial frame rather than a random draw
Answered 200 with an HTML content type1,125Of the rest, 213 answered 403 and 34 answered 503
Disallowed this scanner in robots.txt10Dropped without reading the page
Pages read and analysed1,115The denominator for every figure below
Words of body text in those pages1,269,054Script, style, template, noscript and SVG subtrees excluded
Words inside a marked element121,3129.56 percent of the text read
Pages carrying at least one such word62355.9 percent of the 1,115
Outermost marked regions19,802Nested markers counted once, at the outermost element
One GET of https://<host>/robots.txt and one of https://<host>/ for each hostname, sent as LantadBot/1.0 (+https://lantad.co/bot) with redirects followed and a twenty second timeout, from one network location, with no JavaScript executed and no stylesheet fetched. Where the apex did not answer with HTML the www subdomain was tried once. Measured by Lantad on 30 September 2026 across the 1,419 hostnames in worker/seeds/corpus-seeds-platform.json and worker/seeds/corpus-seeds-industry.json.

Do AI crawlers read hidden text?

Nothing published by any crawler operator answers this, and that absence is the finding rather than a gap in the research. An earlier run read the nine vendor documentation pages behind the fifteen crawler tokens this scanner evaluates and found that two of the nine state whether their crawler executes JavaScript. Not one of them states whether it fetches stylesheets, evaluates a cascade, or drops the contents of an element that carries the HTML hidden attribute. So the honest position is that an extractor with no CSS engine reads every word counted below, an extractor with one reads fewer, and which of those any given engine runs is undocumented.

What can be stated without a vendor's help is what the markers mean, because they are specified. Adding aria-hidden with a value of true removes that element and all of its children from the accessibility tree, which MDN's reference for the attribute, last modified 30 October 2025, states plainly, along with the fact that the attribute hides content from assistive technology without visually hiding anything. The attribute itself is defined in WAI-ARIA, a W3C Recommendation dated 6 June 2023. So a word inside aria-hidden is a word a sighted visitor can see and a screen reader user cannot, and the reason that matters to a measurement of AI visibility is that it is also a word the bytes still carry.

There is one reading of these figures that would be wrong, and it is worth closing off before the numbers start. Google's spam policies page, carrying Last updated 2026-08-28 UTC, defines hidden text or link abuse as the practice of placing content on a page in a way solely to manipulate search engines and not to be easily viewable by human visitors, and gives as examples white text on a white background, text behind an image, CSS positioning off screen, and a font size or opacity of zero. Almost nothing counted in this run looks like that. The dominant pattern is a menu that exists twice so one copy can serve a narrow viewport, and a label that exists so a screen reader has something to announce. This is not a post about spam, and an earlier finding that planted hidden instructions were the weakest of five attacks on AI search agents already covers the adversarial case. This is about drift: the page a machine reads is not the page a person reads, and nobody is checking the gap. What the scanner does and does not count in a grade is set out in the scoring methodology.

  • aria-hidden removes an element from the accessibility tree Stated by MDN's aria-hidden reference, last modified 30 October 2025, and defined in WAI-ARIA, a W3C Recommendation dated 6 June 2023.
  • CSS display can override the HTML hidden attribute Stated by MDN's hidden attribute reference, last modified 17 April 2026: an element styled display block will be displayed despite the attribute.
  • Google treats some hidden text as spam Google's spam policies page, Last updated 2026-08-28 UTC, defines the abuse as content placed solely to manipulate search engines and not easily viewable by visitors.
  • An AI crawler fetches stylesheets No operator documentation behind the fifteen tokens this scanner evaluates says so either way. Undocumented, not disproved.
  • An AI crawler drops content carrying the hidden attribute Same silence. An extractor working from raw bytes has no reason to drop it, and none of the nine vendor pages addresses the question.
  • This run observed a crawler reading these words It did not. One request per host from one network location, no server logs held, no retrieval pipeline inspected.
What is and is not documented about machine consumption of visually hidden markup, as read at the named sources on 30 September 2026. Present means a published page commits to the behaviour, not that the behaviour carries any particular weight.

Four markers, and they hide from different people

Four things in delivered HTML mark an element as not meant for somebody, and all four are readable from the bytes alone. The HTML hidden attribute is the plainest: 232 of the 1,115 pages carried it on an element holding text, 918 such elements in all, 19,724 words between them. An inline style declaring display none appeared on 244 pages, 1,842 elements, 20,020 words. An inline style declaring visibility hidden appeared on 39 pages, 310 elements, 6,838 words. And a class name from the conventional screen reader only set, names like sr-only, visually-hidden and screen-reader-text, appeared on 403 pages, 4,637 elements, 12,476 words. aria-hidden with a value of true is the fifth and by far the largest: 575 pages, 16,512 elements, 65,977 words.

The first four hide from a sighted visitor and leave the text available to assistive technology and to any extractor reading the bytes. aria-hidden does the opposite. 528 pages carried at least one word in the first group, 333 carried at least one under aria-hidden, and 238 carried both, which is the case worth naming: on those 238 pages there are three different readings of the same document, and no two of them agree. That is the same class of problem as a prose parity gap between the crawler view and the browser view, except that here nothing needs to run for the divergence to exist. It is in the bytes as shipped.

One of the four is not what it appears to be, and the distinction decides how much weight the figures carry. The hidden attribute, the two inline styles and aria-hidden are exact: they are either in the markup or they are not. A class name is a convention. This run did not fetch a single stylesheet, so a class named sr-only is evidence that the author intended the element to be hidden visually, not proof that any rule hides it. It is also worth being precise about the hidden attribute, because it is weaker than it looks. MDN's reference for it, last modified 17 April 2026, states that changing the value of the CSS display property on a hidden element will override the hidden state, and that an element styled display block will be displayed despite the attribute's presence. So the attribute is a user agent default, not a directive, and an author rule beats it. Zero of the 19,802 regions used the until-found value, the one form of the attribute that is meant to stay findable.

A second exclusion is worth stating because it moves the numbers. Text inside script, style, template, noscript, SVG, iframe, canvas and select subtrees is not counted anywhere in this post, as words or as marked words. The template case is the one with history: declarative shadow DOM ships component text inside a template element, and a previous note recorded that text inside shadow DOM reaches the browser and not the extractor. Counting a hidden element that sits inside a template would have double counted a subtree this method already treats as absent, which is how an early version of this analysis reported more hidden words on a page than the page had.

MarkerPagesElementsWordsWho cannot reach itExact in the bytes
aria-hidden="true"57516,51265,977Assistive technologyYes
Inline style display none2441,84220,020A sighted visitorYes
HTML hidden attribute23291819,724A sighted visitor, unless CSS overridesYes
Screen reader only class name4034,63712,476A sighted visitor, by conventionNo, no stylesheet fetched
Inline style visibility hidden393106,838A sighted visitorYes
hidden="until-found"000Nobody searching the pageYes
The five markers counted, with who cannot reach the text and how exact the detection is. Pages are of the 1,115 read; an element carrying two markers is counted under both, so the word columns sum to more than the 121,312 total. Measured by Lantad on 30 September 2026.

77 of 820 home pages put every h1 behind a marker

820 of the 1,115 pages declared at least one h1 element. On 77 of those 820, every h1 on the page sat inside an element marked hidden from a sighted visitor. The heading the page nominates as its own subject is therefore present for an extractor and for a screen reader, and absent from the rendered page, which uses a picture, a logo or a styled block of display type instead.

The institutions doing this are not obscure. harvard.edu declares an h1 of Harvard University, mit.edu declares Massachusetts Institute of Technology, stanford.edu declares Stanford University and caltech.edu declares Caltech Homepage, each inside an element carrying a screen reader only class name. time.com declares a pipe separated string of section names. usps.com declares USPS.com Home Page, seattle.gov declares Home Page and sec.gov declares the single word Home. becu.org, a credit union, declares Committed to Your Financial Well-Being on an h1 with an sr-only class, which is the only line on that page making a claim rather than naming a destination. The pattern across the 77 is consistent: where the visible design carries a wordmark, the h1 is supplied for accessibility and it says what the site is called, not what the page is about.

Eleven of the 77 used a marker exact in the bytes rather than a class name, and those are the cases where an extractor honouring CSS and an extractor ignoring it disagree on whether the page has a heading at all. gla.ac.uk, japantimes.co.jp, eltiempo.com, notebooktherapy.com, gldn.com and getalembic.com are among them. Two are worth quoting because the hidden h1 is real prose rather than a label: tailscale.com declares The best secure connectivity platform for the AI era and redis.io declares Inquiring agents want to know, both inside an element that a byte level extractor reads and a rendering one may not. The remaining 66 used a screen reader only class name alone, which this run did not verify against a stylesheet.

This bears on extraction in a specific way rather than a general one. An earlier measurement found that 311 of 704 article pages ran over 300 words with no heading, so a missing heading is already the common case on this corpus. What is new here is a heading that exists and says the wrong thing: an h1 of Home is a navigational label standing in the slot where a machine looks for the page's claim. Put next to the finding that 38 of 1,083 home pages carried no usable name at all, and the finding that 370 of 948 pages left out an Open Graph property the specification requires, the shape is the same each time. The page answers the question what are you more than once, in more than one place, and the answers do not match. That is the mechanism behind weak entity confidence: not an absence of signal, but signals that cannot all be true.

HostnameStratumMarkerh1 as returned
harvard.edueducationsr-only classHarvard University
mit.edueducationsr-only classMassachusetts Institute of Technology
stanford.edueducationsr-only classStanford University
sec.govgovernmentsr-only classHome
seattle.govgovernmentsr-only classHome Page
usps.comgovernmentsr-only classUSPS.com Home Page
time.comnewssr-only classTIME | Current & Breaking News | National & World Updates
becu.orgfinancesr-only classCommitted to Your Financial Well-Being
tailscale.comsaasExact in the bytesThe best secure connectivity platform for the AI era
redis.iosaasExact in the bytesInquiring agents want to know:
japantimes.co.jpnewsExact in the bytesHome page
gla.ac.ukeducationExact in the bytesUniversity of Glasgow
Twelve of the 77 home pages where every h1 sat inside an element marked hidden from a sighted visitor, with the h1 text quoted as returned. Marker is the one carried by the h1 or its nearest marked ancestor. Measured by Lantad on 30 September 2026.

The largest single marked region was a JSON bundle in a textarea

alibaba.com puts 2,054 words inside two textarea elements, both carrying an inline style of display none, with the ids pageHeaderData and pageInitData. Their contents are not prose at all. They are JSON translation and configuration bundles, and they open with strings like headerI18n and header_signin_93 next to human readable values such as Sign in to view message details and Get logistics quotes customized to your needs. A word counter cannot tell the difference between that and a paragraph, and neither can an extractor that flattens the bytes: 2,054 of the page's 2,293 body words are this material, which is why alibaba.com reads as 92 percent hidden.

This is the same failure mode as a hydration payload, already measured here twice. Two golden fixtures in this repository score an identical composite whether the words ship as prose or inside a JSON payload, and a rendering comparison found that 11 of 271 pages returned nothing readable at all until the bundle ran. The difference is the container. A script tag is excluded by every extractor worth the name, including this one, and a textarea is not: it is a form control that legitimately holds user text, so dropping its contents would lose real content elsewhere. Putting a configuration blob in one and hiding it with CSS is therefore invisible to a person, invisible to a screen reader, and fully visible to anything reading the response body.

The screen reader only class names tell a quieter version of the same story, and they are worth reading as a list because the list is short and repetitive. Across the 4,637 elements carrying such a class, the most common contents were the single word image, on 141 of them, then regular price on 102, search on 88, Stat Plus on 67, open dropdown menu on 66, like on 63, close menu on 62, open menu on 61, skip to content on 52, unit price on 52 and skip to main content on 51. Those are interface labels, and three of them, regular price, unit price and sale price, are the default accessibility strings a Shopify theme emits next to a number. They are correct accessibility practice. They are also 4,637 fragments of text that an extractor reads as page content, sitting alongside the real prices, which is a small and systematic distortion of what a product page appears to say.

Where a marked region holds something an answer engine would actually want, the shape is usually tabular. 1,301 of the 19,802 regions contained at least one link with an href, and between them those regions carried 71,495 words, which is 58.9 percent of all the marked text on a fifteenth of the regions. That concentration is the practical summary of this section: most marked regions are tiny labels, and the few large ones are lists and grids. An earlier count found 87 tables on 1,079 home pages, of which 2 carried a caption, so a marked grid of values is not competing with a lot of well formed alternatives for the same facts.

Text inside the elementElementsWhat it is
image141A label for an image link or figure
regular price102A Shopify theme price label
search88A label for a search input or icon button
stat plus:67A subscription tier prefix on article links
open dropdown menu66A control label for a navigation toggle
like63A control label for a reaction button
close menu62The paired label for the same toggle
open menu61The other paired label
skip to content52The skip link, the oldest use of this pattern
unit price52A Shopify theme price label
skip to main content51The same skip link, worded differently
sale price51A Shopify theme price label
share article47A control label on a share button
The most common text found inside elements carrying a conventional screen reader only class name, counted across the 4,637 such elements on 403 of the 1,115 pages read by Lantad on 30 September 2026. Contents are lowercased and truncated to forty characters before counting, so near duplicates group together.

Which stacks and sectors ship the most marked text

The corpus is stratified two ways, by the platform a site is built on and by the sector it operates in, and the platform split is the sharper of the two. The 31 Framer pages that answered put 8,443 of their 39,473 words behind a marker, 21.4 percent, the highest of any stratum by a wide margin and consistent with what a visual site builder produces: carousels, tabbed panels and hover states, all shipped in the markup with only one state visible. At the other end, the 31 static documentation pages put 683 of 29,138 words behind a marker, 2.3 percent, and the 42 Webflow pages put 1,106 of 43,932, 2.5 percent.

That spread is not a quality judgement and it should not be read as one. A documentation site has little interface to label and no product carousel to hide, so it has little reason to mark anything. A travel or ecommerce home page is mostly interface. The 81 travel pages put 14.6 percent behind a marker and the 68 ecommerce pages 12.2 percent, and both numbers are what a large faceted navigation looks like when it is shipped twice for two viewports. The interesting comparison is within a kind: the 34 Shopify pages put 4.2 percent behind a marker while the 68 general ecommerce pages put 12.2 percent, which says the theme layer is doing something more disciplined than hand assembled ecommerce markup does.

Two strata deserve a note because their content is public service rather than commerce. The 101 government pages put 8,350 of 68,697 words behind a marker and the 101 healthcare pages 14,013 of 107,059, the second highest industry share. tewhatuora.govt.nz, the New Zealand health authority, is the largest single case in the corpus outside Framer: 3,591 of its 4,831 words, including one div carrying both the hidden attribute and aria-hidden and holding 1,282 words of condition names, from Allergies through Bones, muscles and joints. That is a reference index of exactly the kind an answer engine would want to quote, marked as not meant for two of its three audiences.

None of this is scored. Lantad does not weight marked text in a grade, this post is not an argument that it should, and a site with a high share here is not a site with a problem. The claim is narrower. The words a machine extracts from these pages are not the words on these pages, the gap reaches 30 percent at the ninetieth percentile, and no engine publishes which side of it reads. If you want to see which of your own words survive the trip, the GPTBot view of a page shows the extracted text, and the corpus level figures behind posts like this one are collected on the research page. The broader question of what actually moves AI visibility is not settled by a word count, and this post does not pretend otherwise.

StratumFramePagesBody wordsMarked wordsShare
framerplatform3139,4738,44321.4 percent
travelindustry8191,32913,29714.6 percent
healthcareindustry101107,05914,01313.1 percent
saasindustry117182,03323,74413.0 percent
ecommerceindustry6883,09510,11512.2 percent
governmentindustry10168,6978,35012.2 percent
spa-startupsplatform4447,8854,76710.0 percent
bubble-nocodeplatform4234,2553,3169.7 percent
financeindustry103130,23111,6569.0 percent
educationindustry105107,0189,4168.8 percent
saas-marketingplatform3746,2572,6555.7 percent
shopify-dtcplatform3442,4061,7704.2 percent
wordpress-smbplatform3640,2771,5803.9 percent
newsindustry58119,2504,6703.9 percent
wix-squarespaceplatform5431,4911,0313.3 percent
media-localplatform3025,2287002.8 percent
webflowplatform4243,9321,1062.5 percent
static-docsplatform3129,1386832.3 percent
Share of body words sitting inside a marked element, by corpus stratum. Pages is the number in that stratum that answered 200 with HTML and was not disallowed in robots.txt. Measured by Lantad on 30 September 2026.

What this run did not measure

No stylesheet was fetched and no cascade was evaluated, which is the largest limit and it cuts both ways. An element hidden by an external rule, a media query or a class this run does not recognise is counted as visible here, so 121,312 is a floor rather than an estimate. And a screen reader only class name is an authorial intention rather than an observed effect, which is why the 12,476 words behind such class names are reported separately from the 46,582 behind the three markers that hide from a sighted visitor and are exact in the bytes. Anyone repeating this measurement with a browser would get different and larger figures, and the two methods answer different questions: this one asks what the response body says, which is what a byte level extractor gets.

No JavaScript ran, so an element revealed or marked by script after load is recorded in whatever state the server sent. That matters more on some stacks than others: a Next.js site may ship every tab panel in the initial payload and reveal one on hydration, and this method sees all of them. One page was requested per site, always the home page, and a home page is the most interface heavy page a site has, so these shares are almost certainly above what the same sites' article and product pages would show. Requests came from one network location on one date, and 294 of the 1,419 hostnames did not answer with HTML at all, 213 of them with a 403, so the readable set is not a random sample of the corpus and the corpus is not a random sample of the web.

Most importantly, no crawler was observed doing anything. This run fetched pages; it did not inspect any engine's retrieval pipeline, hold any server logs, or test whether a marked word ever reached a model. The honest statement of the finding is a statement about documents: 9.56 percent of the words these 1,115 pages served on 30 September 2026 were marked as not meant for at least one of the page's audiences, the share reaches 30.5 percent at the ninetieth percentile, and on 77 of the 820 pages with an h1 the page's own heading was among them. What any given extractor does with that text is undocumented by every operator whose token this scanner evaluates, and undocumented is the result rather than a gap in the method.

  • Markers in the bytes Measured The hidden attribute, inline display none, inline visibility hidden and aria-hidden are exact in the markup. 46,582 words sat behind one of the first three and 65,977 behind aria-hidden.
  • Screen reader only classes Convention only 403 pages, 12,476 words, detected by class name. No stylesheet was fetched, so no rule was confirmed.
  • External CSS Not measured No stylesheet fetched and no cascade evaluated, so an element hidden by an external rule counts as visible and 121,312 is a floor.
  • Rendered state Not measured No JavaScript executed. An element revealed or marked after load is recorded as the server sent it.
  • Crawler behaviour Undocumented No operator behind the fifteen tokens this scanner evaluates publishes whether its extractor fetches CSS or drops marked content.
  • Interior pages Out of scope One request per site, always the home page, which is the most interface heavy page a site has.
What this measurement establishes and what it does not, stated against the method actually run on 30 September 2026.

Written by

Lantad

Published .

A page has more than one audience, and they do not all get the same text. A sighted visitor reads what the stylesheet lets through. Assistive technology reads what the accessibility tree holds. An extractor that takes the delivered bytes and pulls the words out of them reads whatever is between the tags, because it has no viewport and, in most cases, no CSS engine. Those three readings can differ, and on a normal commercial home page they do.

Common questions

Do AI crawlers read hidden text?

Nobody publishes an answer. None of the nine vendor documentation pages behind the fifteen crawler tokens Lantad evaluates states whether its crawler fetches stylesheets or drops content carrying the HTML hidden attribute. What can be said is that an extractor working from the response body has no mechanism for excluding it: on the 1,115 home pages read on 30 September 2026, 121,312 of 1,269,054 body words sat inside an element marked hidden from a sighted visitor or from assistive technology.

Is hidden text on my site a Google spam problem?

Almost certainly not, on the evidence of this corpus. Google's spam policies page, Last updated 2026-08-28 UTC, defines the abuse as content placed solely to manipulate search engines and not easily viewable by visitors, and gives white text on a white background and off screen positioning as examples. What this run found instead was duplicated navigation for narrow viewports and accessibility labels, which is neither of those things.

Does the HTML hidden attribute stop a crawler reading an element?

There is no published commitment that it stops any crawler, and it does not reliably stop a browser either. MDN's reference for the attribute, last modified 17 April 2026, states that changing the CSS display property on a hidden element overrides the hidden state. So the attribute is a user agent default rather than a directive, and 232 of the 1,115 pages read here carried it on an element holding text, 19,724 words in all.

Should an h1 be hidden with a screen reader only class?

That is an accessibility decision rather than a visibility one, and this measurement does not argue against it. What it records is the consequence for extraction: on 77 of the 820 pages that declared an h1, every h1 sat inside a marked element, and the text was usually the site's name or the word Home rather than a claim about the page. A machine looking for what the page is about finds a navigational label in the slot reserved for the answer.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.