BlogFindings
Do AI crawlers read hidden text? 121,312 of 1,269,054 words sat behind a marker on 1,115 home pages
Lantad requested the robots.txt and then the home page of all 1,419 hostnames in this repository's committed corpus on 30 September 2026 and read the delivered bytes with no JavaScript executed and no stylesheet fetched. 1,115 pages were both allowed and readable, and they carried 1,269,054 words of body text. 121,312 of those words, 9.56 percent, sat inside an element marked hidden from somebody: from a sighted visitor, or from assistive technology, or from both. On 77 of the 820 pages that declared an h1, every h1 was inside one of those elements.
This run counted the difference. Lantad requested the robots.txt of all 1,419 hostnames in this repository's two committed corpus seed files on 30 September 2026, then the home page of each, sent as its own declared crawler user agent with redirects followed and a twenty second timeout, from one network location, with no JavaScript executed and no stylesheet fetched. 1,125 answered HTTP 200 with an HTML content type. 12 of the 1,419 robots.txt files disallowed this scanner at the site root, and 10 of those 12 were among the 1,125, so they were dropped rather than read. That leaves 1,115 pages, carrying 1,269,054 words of body text, and the question is how many of those words an AI crawler would find that a person would not.
In short
- Do AI crawlers read hidden text is answerable in one direction only: Lantad read 1,115 home pages on 30 September 2026 and found 121,312 of their 1,269,054 body words inside an element marked hidden from somebody, but no crawler operator publishes what its extractor does with any of the four markers counted here.
- 623 of the 1,115 pages carried at least one such word on 30 September 2026, and the marker was not evenly used: aria-hidden held 65,977 words across 575 pages, an inline display none style held 20,020 across 244, the HTML hidden attribute held 19,724 across 232, and a screen reader only class name held 12,476 across 403.
- 77 of the 820 home pages that declared an h1 put every one of them inside a marked element, harvard.edu, mit.edu, time.com and usps.com among them, and 66 of those 77 used a screen reader only class name rather than a marker that is exact in the delivered bytes.
- 12,393 of the 19,802 marked regions sat outside any nav, header, footer or aside element and outside any navigation, banner, contentinfo, menubar or menu role, carrying 74,897 words, so the extra text is not all duplicated menus even though most of the largest single regions are.
- The Framer stratum carried the highest share at 8,443 of 39,473 words on 31 pages, and the static documentation stratum the lowest at 683 of 29,138 on 31 pages, measured by Lantad on 30 September 2026.
| Stage | Count | What happened |
|---|---|---|
| Hostnames requested | 1,419 | The committed corpus, an editorial frame rather than a random draw |
| Answered 200 with an HTML content type | 1,125 | Of the rest, 213 answered 403 and 34 answered 503 |
| Disallowed this scanner in robots.txt | 10 | Dropped without reading the page |
| Pages read and analysed | 1,115 | The denominator for every figure below |
| Words of body text in those pages | 1,269,054 | Script, style, template, noscript and SVG subtrees excluded |
| Words inside a marked element | 121,312 | 9.56 percent of the text read |
| Pages carrying at least one such word | 623 | 55.9 percent of the 1,115 |
| Outermost marked regions | 19,802 | Nested markers counted once, at the outermost element |
Four markers, and they hide from different people
Four things in delivered HTML mark an element as not meant for somebody, and all four are readable from the bytes alone. The HTML hidden attribute is the plainest: 232 of the 1,115 pages carried it on an element holding text, 918 such elements in all, 19,724 words between them. An inline style declaring display none appeared on 244 pages, 1,842 elements, 20,020 words. An inline style declaring visibility hidden appeared on 39 pages, 310 elements, 6,838 words. And a class name from the conventional screen reader only set, names like sr-only, visually-hidden and screen-reader-text, appeared on 403 pages, 4,637 elements, 12,476 words. aria-hidden with a value of true is the fifth and by far the largest: 575 pages, 16,512 elements, 65,977 words.
The first four hide from a sighted visitor and leave the text available to assistive technology and to any extractor reading the bytes. aria-hidden does the opposite. 528 pages carried at least one word in the first group, 333 carried at least one under aria-hidden, and 238 carried both, which is the case worth naming: on those 238 pages there are three different readings of the same document, and no two of them agree. That is the same class of problem as a prose parity gap between the crawler view and the browser view, except that here nothing needs to run for the divergence to exist. It is in the bytes as shipped.
One of the four is not what it appears to be, and the distinction decides how much weight the figures carry. The hidden attribute, the two inline styles and aria-hidden are exact: they are either in the markup or they are not. A class name is a convention. This run did not fetch a single stylesheet, so a class named sr-only is evidence that the author intended the element to be hidden visually, not proof that any rule hides it. It is also worth being precise about the hidden attribute, because it is weaker than it looks. MDN's reference for it, last modified 17 April 2026, states that changing the value of the CSS display property on a hidden element will override the hidden state, and that an element styled display block will be displayed despite the attribute's presence. So the attribute is a user agent default, not a directive, and an author rule beats it. Zero of the 19,802 regions used the until-found value, the one form of the attribute that is meant to stay findable.
A second exclusion is worth stating because it moves the numbers. Text inside script, style, template, noscript, SVG, iframe, canvas and select subtrees is not counted anywhere in this post, as words or as marked words. The template case is the one with history: declarative shadow DOM ships component text inside a template element, and a previous note recorded that text inside shadow DOM reaches the browser and not the extractor. Counting a hidden element that sits inside a template would have double counted a subtree this method already treats as absent, which is how an early version of this analysis reported more hidden words on a page than the page had.
| Marker | Pages | Elements | Words | Who cannot reach it | Exact in the bytes |
|---|---|---|---|---|---|
| aria-hidden="true" | 575 | 16,512 | 65,977 | Assistive technology | Yes |
| Inline style display none | 244 | 1,842 | 20,020 | A sighted visitor | Yes |
| HTML hidden attribute | 232 | 918 | 19,724 | A sighted visitor, unless CSS overrides | Yes |
| Screen reader only class name | 403 | 4,637 | 12,476 | A sighted visitor, by convention | No, no stylesheet fetched |
| Inline style visibility hidden | 39 | 310 | 6,838 | A sighted visitor | Yes |
| hidden="until-found" | 0 | 0 | 0 | Nobody searching the page | Yes |
77 of 820 home pages put every h1 behind a marker
820 of the 1,115 pages declared at least one h1 element. On 77 of those 820, every h1 on the page sat inside an element marked hidden from a sighted visitor. The heading the page nominates as its own subject is therefore present for an extractor and for a screen reader, and absent from the rendered page, which uses a picture, a logo or a styled block of display type instead.
The institutions doing this are not obscure. harvard.edu declares an h1 of Harvard University, mit.edu declares Massachusetts Institute of Technology, stanford.edu declares Stanford University and caltech.edu declares Caltech Homepage, each inside an element carrying a screen reader only class name. time.com declares a pipe separated string of section names. usps.com declares USPS.com Home Page, seattle.gov declares Home Page and sec.gov declares the single word Home. becu.org, a credit union, declares Committed to Your Financial Well-Being on an h1 with an sr-only class, which is the only line on that page making a claim rather than naming a destination. The pattern across the 77 is consistent: where the visible design carries a wordmark, the h1 is supplied for accessibility and it says what the site is called, not what the page is about.
Eleven of the 77 used a marker exact in the bytes rather than a class name, and those are the cases where an extractor honouring CSS and an extractor ignoring it disagree on whether the page has a heading at all. gla.ac.uk, japantimes.co.jp, eltiempo.com, notebooktherapy.com, gldn.com and getalembic.com are among them. Two are worth quoting because the hidden h1 is real prose rather than a label: tailscale.com declares The best secure connectivity platform for the AI era and redis.io declares Inquiring agents want to know, both inside an element that a byte level extractor reads and a rendering one may not. The remaining 66 used a screen reader only class name alone, which this run did not verify against a stylesheet.
This bears on extraction in a specific way rather than a general one. An earlier measurement found that 311 of 704 article pages ran over 300 words with no heading, so a missing heading is already the common case on this corpus. What is new here is a heading that exists and says the wrong thing: an h1 of Home is a navigational label standing in the slot where a machine looks for the page's claim. Put next to the finding that 38 of 1,083 home pages carried no usable name at all, and the finding that 370 of 948 pages left out an Open Graph property the specification requires, the shape is the same each time. The page answers the question what are you more than once, in more than one place, and the answers do not match. That is the mechanism behind weak entity confidence: not an absence of signal, but signals that cannot all be true.
| Hostname | Stratum | Marker | h1 as returned |
|---|---|---|---|
| harvard.edu | education | sr-only class | Harvard University |
| mit.edu | education | sr-only class | Massachusetts Institute of Technology |
| stanford.edu | education | sr-only class | Stanford University |
| sec.gov | government | sr-only class | Home |
| seattle.gov | government | sr-only class | Home Page |
| usps.com | government | sr-only class | USPS.com Home Page |
| time.com | news | sr-only class | TIME | Current & Breaking News | National & World Updates |
| becu.org | finance | sr-only class | Committed to Your Financial Well-Being |
| tailscale.com | saas | Exact in the bytes | The best secure connectivity platform for the AI era |
| redis.io | saas | Exact in the bytes | Inquiring agents want to know: |
| japantimes.co.jp | news | Exact in the bytes | Home page |
| gla.ac.uk | education | Exact in the bytes | University of Glasgow |
Which stacks and sectors ship the most marked text
The corpus is stratified two ways, by the platform a site is built on and by the sector it operates in, and the platform split is the sharper of the two. The 31 Framer pages that answered put 8,443 of their 39,473 words behind a marker, 21.4 percent, the highest of any stratum by a wide margin and consistent with what a visual site builder produces: carousels, tabbed panels and hover states, all shipped in the markup with only one state visible. At the other end, the 31 static documentation pages put 683 of 29,138 words behind a marker, 2.3 percent, and the 42 Webflow pages put 1,106 of 43,932, 2.5 percent.
That spread is not a quality judgement and it should not be read as one. A documentation site has little interface to label and no product carousel to hide, so it has little reason to mark anything. A travel or ecommerce home page is mostly interface. The 81 travel pages put 14.6 percent behind a marker and the 68 ecommerce pages 12.2 percent, and both numbers are what a large faceted navigation looks like when it is shipped twice for two viewports. The interesting comparison is within a kind: the 34 Shopify pages put 4.2 percent behind a marker while the 68 general ecommerce pages put 12.2 percent, which says the theme layer is doing something more disciplined than hand assembled ecommerce markup does.
Two strata deserve a note because their content is public service rather than commerce. The 101 government pages put 8,350 of 68,697 words behind a marker and the 101 healthcare pages 14,013 of 107,059, the second highest industry share. tewhatuora.govt.nz, the New Zealand health authority, is the largest single case in the corpus outside Framer: 3,591 of its 4,831 words, including one div carrying both the hidden attribute and aria-hidden and holding 1,282 words of condition names, from Allergies through Bones, muscles and joints. That is a reference index of exactly the kind an answer engine would want to quote, marked as not meant for two of its three audiences.
None of this is scored. Lantad does not weight marked text in a grade, this post is not an argument that it should, and a site with a high share here is not a site with a problem. The claim is narrower. The words a machine extracts from these pages are not the words on these pages, the gap reaches 30 percent at the ninetieth percentile, and no engine publishes which side of it reads. If you want to see which of your own words survive the trip, the GPTBot view of a page shows the extracted text, and the corpus level figures behind posts like this one are collected on the research page. The broader question of what actually moves AI visibility is not settled by a word count, and this post does not pretend otherwise.
| Stratum | Frame | Pages | Body words | Marked words | Share |
|---|---|---|---|---|---|
| framer | platform | 31 | 39,473 | 8,443 | 21.4 percent |
| travel | industry | 81 | 91,329 | 13,297 | 14.6 percent |
| healthcare | industry | 101 | 107,059 | 14,013 | 13.1 percent |
| saas | industry | 117 | 182,033 | 23,744 | 13.0 percent |
| ecommerce | industry | 68 | 83,095 | 10,115 | 12.2 percent |
| government | industry | 101 | 68,697 | 8,350 | 12.2 percent |
| spa-startups | platform | 44 | 47,885 | 4,767 | 10.0 percent |
| bubble-nocode | platform | 42 | 34,255 | 3,316 | 9.7 percent |
| finance | industry | 103 | 130,231 | 11,656 | 9.0 percent |
| education | industry | 105 | 107,018 | 9,416 | 8.8 percent |
| saas-marketing | platform | 37 | 46,257 | 2,655 | 5.7 percent |
| shopify-dtc | platform | 34 | 42,406 | 1,770 | 4.2 percent |
| wordpress-smb | platform | 36 | 40,277 | 1,580 | 3.9 percent |
| news | industry | 58 | 119,250 | 4,670 | 3.9 percent |
| wix-squarespace | platform | 54 | 31,491 | 1,031 | 3.3 percent |
| media-local | platform | 30 | 25,228 | 700 | 2.8 percent |
| webflow | platform | 42 | 43,932 | 1,106 | 2.5 percent |
| static-docs | platform | 31 | 29,138 | 683 | 2.3 percent |
What this run did not measure
No stylesheet was fetched and no cascade was evaluated, which is the largest limit and it cuts both ways. An element hidden by an external rule, a media query or a class this run does not recognise is counted as visible here, so 121,312 is a floor rather than an estimate. And a screen reader only class name is an authorial intention rather than an observed effect, which is why the 12,476 words behind such class names are reported separately from the 46,582 behind the three markers that hide from a sighted visitor and are exact in the bytes. Anyone repeating this measurement with a browser would get different and larger figures, and the two methods answer different questions: this one asks what the response body says, which is what a byte level extractor gets.
No JavaScript ran, so an element revealed or marked by script after load is recorded in whatever state the server sent. That matters more on some stacks than others: a Next.js site may ship every tab panel in the initial payload and reveal one on hydration, and this method sees all of them. One page was requested per site, always the home page, and a home page is the most interface heavy page a site has, so these shares are almost certainly above what the same sites' article and product pages would show. Requests came from one network location on one date, and 294 of the 1,419 hostnames did not answer with HTML at all, 213 of them with a 403, so the readable set is not a random sample of the corpus and the corpus is not a random sample of the web.
Most importantly, no crawler was observed doing anything. This run fetched pages; it did not inspect any engine's retrieval pipeline, hold any server logs, or test whether a marked word ever reached a model. The honest statement of the finding is a statement about documents: 9.56 percent of the words these 1,115 pages served on 30 September 2026 were marked as not meant for at least one of the page's audiences, the share reaches 30.5 percent at the ninetieth percentile, and on 77 of the 820 pages with an h1 the page's own heading was among them. What any given extractor does with that text is undocumented by every operator whose token this scanner evaluates, and undocumented is the result rather than a gap in the method.
-
Markers in the bytesMeasured The hidden attribute, inline display none, inline visibility hidden and aria-hidden are exact in the markup. 46,582 words sat behind one of the first three and 65,977 behind aria-hidden. -
Screen reader only classesConvention only 403 pages, 12,476 words, detected by class name. No stylesheet was fetched, so no rule was confirmed. -
External CSSNot measured No stylesheet fetched and no cascade evaluated, so an element hidden by an external rule counts as visible and 121,312 is a floor. -
Rendered stateNot measured No JavaScript executed. An element revealed or marked after load is recorded as the server sent it. -
Crawler behaviourUndocumented No operator behind the fifteen tokens this scanner evaluates publishes whether its extractor fetches CSS or drops marked content. -
Interior pagesOut of scope One request per site, always the home page, which is the most interface heavy page a site has.
Lantad
Published .
A page has more than one audience, and they do not all get the same text. A sighted visitor reads what the stylesheet lets through. Assistive technology reads what the accessibility tree holds. An extractor that takes the delivered bytes and pulls the words out of them reads whatever is between the tags, because it has no viewport and, in most cases, no CSS engine. Those three readings can differ, and on a normal commercial home page they do.
Common questions
Do AI crawlers read hidden text?
Nobody publishes an answer. None of the nine vendor documentation pages behind the fifteen crawler tokens Lantad evaluates states whether its crawler fetches stylesheets or drops content carrying the HTML hidden attribute. What can be said is that an extractor working from the response body has no mechanism for excluding it: on the 1,115 home pages read on 30 September 2026, 121,312 of 1,269,054 body words sat inside an element marked hidden from a sighted visitor or from assistive technology.
Is hidden text on my site a Google spam problem?
Almost certainly not, on the evidence of this corpus. Google's spam policies page, Last updated 2026-08-28 UTC, defines the abuse as content placed solely to manipulate search engines and not easily viewable by visitors, and gives white text on a white background and off screen positioning as examples. What this run found instead was duplicated navigation for narrow viewports and accessibility labels, which is neither of those things.
Does the HTML hidden attribute stop a crawler reading an element?
There is no published commitment that it stops any crawler, and it does not reliably stop a browser either. MDN's reference for the attribute, last modified 17 April 2026, states that changing the CSS display property on a hidden element overrides the hidden state. So the attribute is a user agent default rather than a directive, and 232 of the 1,115 pages read here carried it on an element holding text, 19,724 words in all.
Should an h1 be hidden with a screen reader only class?
That is an accessibility decision rather than a visibility one, and this measurement does not argue against it. What it records is the consequence for extraction: on 77 of the 820 pages that declared an h1, every h1 sat inside a marked element, and the text was usually the site's name or the word Home rather than a claim about the page. A machine looking for what the page is about finds a navigational label in the slot reserved for the answer.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.