BlogFindings
Do AI crawlers read web components? 3,291 of 7,124 custom elements arrived with no text
Lantad requested the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 8 October 2026 as LantadBot, following redirects and executing no JavaScript. 1,077 answered HTTP 200 with HTML. 219 of those carried at least one custom element, 7,124 instances across 1,090 distinct tag names, and 3,291 of those instances arrived with no text inside them.
So we counted. On 8 October 2026 Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for their home page, as LantadBot, following redirects, running no JavaScript, and 1,077 answered HTTP 200 with an HTML content type. 219 of those pages carried at least one custom element. Between them they held 7,124 instances across 1,090 distinct tag names, and 3,291 of the instances contained no text at all in the bytes the server sent. The interesting part is not that number. It is that the same number covers an image wrapper that was never going to hold a sentence and an airline home page whose entire content sat behind a single tag.
In short
- Lantad requested the home page of all 1,419 hostnames in this repository's committed corpus on 8 October 2026 and 1,077 answered HTTP 200 with HTML: 219 of those carried at least one custom element, and 3,291 of the 7,124 instances arrived with no text inside them.
- Do AI crawlers read web components: measured on 8 October 2026, 8 of the 1,077 readable home pages served zero words of body text with the whole page sitting behind one empty custom element, among them ryanair.com, archive.org, bcb.gov.br and buenosaires.gob.ar.
- Declarative shadow DOM, the server rendered form that puts component text into the delivered bytes, appeared on 16 of the 1,077 pages in 520 template elements holding 38,135 words.
- Lantad's own extraction ruleset drops template subtrees, so on stanford.edu on 8 October 2026 this scanner read 1,385 words while 10,436 words sat inside declarative shadow roots it never opened.
- 96 of the 219 pages carrying a custom element had every one of them arrive empty, while on 13 other pages every word of body text sat inside one, so an empty custom element is not on its own a measure of anything lost.
| What was counted | Figure |
|---|---|
| Hostnames requested | 1,419 |
| Answered HTTP 200 with HTML | 1,077 |
| Pages carrying at least one custom element | 219 |
| Custom element instances | 7,124 |
| Instances with no text in the bytes | 3,291 |
| Instances carrying text | 3,833 |
| Distinct tag names | 1,090 |
| Pages using declarative shadow DOM | 16 |
A custom element is a tag the parser does not recognise
The naming rule is what makes these countable. Mozilla's guide to using custom elements, last modified on 1 September 2026, states that the name passed to define must start with a lowercase letter and contain a hyphen. The hyphen is not a convention. It is the reservation that keeps author invented tags from ever colliding with a tag the standard might add later, and it means any lowercase hyphenated tag in a document is either a custom element or one of eight names the specification keeps for itself, among them annotation-xml and font-face. The eight are enumerated in the HTML Standard's definition of a valid custom element name, at html.spec.whatwg.org/multipage/custom-elements.html, which says they are there because SVG 2 and MathML got to the hyphen first. Excluding those eight, a parser can separate the two sets with no knowledge of the site.
That is the method here. The 7,124 instances were found by driving the same streaming parser the scanner uses over the raw response and collecting every lowercase hyphenated tag that is not one of the eight reserved names. Nothing was rendered, nothing was executed, and no browser was involved, which is deliberate: a crawler that does not run JavaScript sees exactly this document and nothing more.
Text can then sit in one of three places. It can sit between the opening and closing tag as ordinary child content, which the standard calls the light DOM, and in that case it is ordinary text that any parser reads without knowing what the component does. It can sit inside a template element carrying a shadowrootmode attribute, which is in the bytes but is not ordinary text. Or it can exist only in JavaScript, in which case the response contains an empty tag and a promise. The first case costs a crawler nothing. The third costs it everything. Which of the three a site chose is not visible from the outside without opening the bytes, and it is not what prose parity measures either, because parity compares a crawler fetch against a rendered one rather than asking where in the document the words were. An AI crawler does not report which of the three it met. It just sends a request and keeps whatever came back.
Sample Illustrative, not a measurement of any real site.
Flow: Server sends HTML to Text as child content; Server sends HTML to Text in a shadowrootmode template; Server sends HTML to Text only in JavaScript; Text as child content (read as ordinary text) to Crawler reads the words; Text in a shadowrootmode template (dropped with the template) to Crawler reads an empty tag; Text only in JavaScript (never arrived) to Crawler reads an empty tag.
Do AI crawlers read web components?
Only one operator documents an answer. Google Search Central's JavaScript SEO basics, carrying Last updated 2026-03-04 UTC, states that Google supports web components, that when Google renders a page it flattens the shadow DOM and light DOM content, that this means Google can only see content that is visible in the rendered HTML, and that if the content is not visible in the rendered HTML then Google will not be able to index it. Read carefully, that is a conditional rather than a reassurance: the support is real and it is support for rendering, so it transfers only to a client that renders.
Most do not say. Lantad read the published crawler documentation of nine AI operators and found that two of nine state either way whether they execute JavaScript, which leaves the component question unanswered for the rest by their own choice. For a page whose text sits in the light DOM this does not matter, because there is nothing to execute. For a page whose text arrives only after a script runs, the operator's silence is the whole risk, and it is the same risk the AI visibility question always reduces to: not what a crawler is capable of, but what your server put in the response it actually received.
There is one more wrinkle that the flattening sentence hides. Flattening describes what a renderer does with a shadow root it has, and a declarative shadow root only becomes a shadow root if the parser that read the document implemented the declarative form. A parser that did not implement it produces an ordinary template element instead, and template content is not rendered. So the same bytes can yield component text to one consumer and nothing to another, with no error anywhere. The cheapest way to see which you are serving is to fetch your own page the way a crawler does, with what GPTBot sees, and read the result rather than the intention.
| Where the text sits | In the response bytes | Read by a parser that runs no JavaScript | What Google documents |
|---|---|---|---|
| Child content of the tag | Yes | Yes | Flattened with the shadow DOM when rendered |
| Inside a shadowrootmode template | Yes | Only if the parser implements the declarative form | Flattened with the light DOM when rendered |
| Assembled by script at runtime | No | No | Not indexable if absent from the rendered HTML |
What 1,077 home pages carried on 8 October 2026
The ladder from 1,419 hostnames to 1,077 pages is worth stating because every rate below rests on it. 11 sites closed their own root to this scanner in robots.txt and were not requested, which is the behaviour the scanner's published bot page promises. 48 never returned a status, 25 of them timing out and 23 failing in the client. 216 answered HTTP 403 and 40 answered HTTP 503, which together are 256 sites that served this identified crawler no page at all. 27 more answered some other status. The remaining 1,077 answered HTTP 200 with an HTML content type and are the denominator for everything that follows, exactly as the methodology page describes for every study published here.
219 of those 1,077 carried a custom element, and the spread by stratum is wider than any other result in this corpus. 26 of 34 readable Shopify direct to consumer storefronts carried one, 76.5 percent, which is what you would expect of a platform whose default theme ships components named cart-drawer, predictive-search and header-drawer: those three alone appeared on 11, 14 and 12 pages. Against that, 2 of 36 WordPress small business pages carried one, and 1 of 37 SaaS marketing pages. The per stack guidance for Shopify is therefore load bearing in a way it is not for a WordPress theme, because a Shopify merchant inherits components whether or not anyone chose them.
The tag names themselves say how little of this is shared vocabulary. Of the 1,090 distinct names, 999 appeared on exactly one page and only 91 appeared on two or more. The most widespread single name was wow-image, a Wix image wrapper, on 19 pages in 278 instances, every one of them empty. Next was astro-island on 16 pages, the hydration wrapper Astro emits, of which 102 of 155 instances were empty. Then wix-dropdown-menu on 15 pages, which was never empty once. The median page carrying any custom element carried five of them, and one page carried 1,612.
| Stratum | Readable | With a custom element | Percent | Instances | Empty |
|---|---|---|---|---|---|
| Shopify direct to consumer | 34 | 26 | 76.5 | 1,276 | 474 |
| Local media | 30 | 11 | 36.7 | 64 | 62 |
| Wix and Squarespace | 54 | 18 | 33.3 | 143 | 125 |
| Travel | 78 | 24 | 30.8 | 403 | 243 |
| News | 60 | 14 | 23.3 | 2,267 | 1,259 |
| Finance | 95 | 21 | 22.1 | 563 | 207 |
| Government | 86 | 18 | 20.9 | 386 | 168 |
| SaaS | 116 | 23 | 19.8 | 874 | 342 |
| Education | 101 | 10 | 9.9 | 74 | 48 |
| Webflow | 42 | 4 | 9.5 | 16 | 14 |
| WordPress small business | 36 | 2 | 5.6 | 25 | 4 |
Eight pages served a crawler zero words with the page behind one tag
The clearest cases in the corpus are the ones where a single empty custom element is the document. archive.org answered this crawler with 1,872 bytes containing a title, a stylesheet link, a script tag loading the Lit polyfill and the webcomponentsjs loader, and an empty app-root element. bcb.gov.br, the Central Bank of Brazil, answered with 2,871 bytes and an empty app-root. ryanair.com answered with 4,743 bytes and an empty hp-app-root. Add aircanada.com with ac-home-root, bloomandwild.com, orcid.org, buenosaires.gob.ar and sundhed.dk, whose three custom elements are named for analytics, a keepalive and a loading throbber, and that is 8 pages serving zero words of body text with their content behind a component that had not run. Two more came close: citi.com served 3 words across 255,867 bytes and 26 custom elements, and visitsingapore.com served 6 words behind 16 tags named stb-masterhead, stb-things-to-do and the like.
The shape recurs because framework defaults produce it. app-root appeared on 10 pages and router-outlet, the Angular element that marks where a route renders, on 8, with all 17 of its instances empty. These are the single page application shells this blog has measured from other directions: the 17 of 380 home pages that sent a crawler zero words and the 11 of 271 pages where JavaScript supplied every word. Our own client rendered canary page exists to hold this exact failure still so a scan can be checked against a known answer, and the React stack guide says what to change.
The honest counterweight belongs here rather than at the end. 55 of the 1,077 readable pages served zero words of body text, and 47 of those 55 carried no custom element at all. Components are therefore a visible cause of an empty page and not the main one. A page can be blank to a crawler for many reasons, and most of the blank pages in this corpus were blank without one.
-
archive.org0 words, 1,872 bytes Empty app-root, with the Lit polyfill and the webcomponentsjs loader in the head. -
ryanair.com0 words, 4,743 bytes Empty hp-app-root and nothing else in the body. -
bcb.gov.br0 words, 2,871 bytes Central Bank of Brazil, empty app-root. -
buenosaires.gob.ar0 words, 37,745 bytes Empty app-root inside a much larger shell of scripts and styles. -
aircanada.com0 words, 71,094 bytes Empty ac-home-root, a renamed root element. -
orcid.org0 words, 67,250 bytes Empty app-root on a scholarly identifier registry. -
bloomandwild.com0 words, 65,218 bytes Empty app-root on a direct to consumer storefront. -
sundhed.dk0 words, 3,235 bytes Three empty elements named for analytics, a keepalive and a throbber.
The server rendered form was on 16 pages, and our own extractor drops it
There is a standard answer to all of this, and almost nobody in the corpus uses it. Mozilla's template element reference, last modified on 30 August 2026, states that if the element carries a shadowrootmode attribute with a value of open or closed then the HTML parser will immediately generate a shadow DOM, and that the element is replaced in the DOM by its content wrapped in a shadow root. The same page states that by default a template element's content is not rendered. Those two sentences are the entire mechanism: the words ship in the response, and whether they become content depends on the parser recognising one attribute. Google's web.dev article on declarative shadow DOM, last updated 2024-08-13, gives the reasons it exists as accessibility and a no JavaScript fallback, and never mentions a search crawler at all.
16 of the 1,077 pages used it, in 520 template elements holding 38,135 words between them. The concentration is extreme: stanford.edu carried 54 such templates holding 10,436 words, collibra.com 114 holding 9,373, redhat.com 135 holding 6,629, tui.com 2 holding 4,738 and johnlewis.com 1 holding 3,283. On 5 of the 16 sites there were more words inside the declarative shadow roots than outside them. Those sites did the harder, better thing: they put their component text in the bytes, where a client that never runs a line of script can reach it. Our own server rendered canary page is the control for what that should look like.
And this scanner does not read any of it. The extraction ruleset in core/src/extract.ts drops script, style, noscript, template, svg and iframe subtrees entirely, template among them, which was the right default for a decade in which template meant inert markup and is now the wrong default for five sites in this corpus. Run over stanford.edu on 8 October 2026, Lantad's extractor returned 1,385 words from a page carrying 10,436 more inside declarative shadow roots. On collibra.com it returned 1,197 against 9,373, and on johnlewis.com 707 against 3,283. We wrote this gap down on 4 August 2026 and recorded in core/src/renderparity.ts that across a random sample of 24 scanned corpus sites no page used declarative shadow DOM, so the effect on the published study was nil, with a note to revisit if a sampled platform turned out to lean on web components. That condition is now met. The sample was 24 sites and the answer was correct for those 24; on 1,077 it is 16 sites and 38,135 words, and the five worst of them would be graded by this scanner on a fraction of their own prose.
| Host | Templates | Words inside shadow roots | Words outside |
|---|---|---|---|
| stanford.edu | 54 | 10,436 | 1,390 |
| collibra.com | 114 | 9,373 | 1,233 |
| redhat.com | 135 | 6,629 | 2,251 |
| tui.com | 2 | 4,738 | 1,342 |
| johnlewis.com | 1 | 3,283 | 719 |
| salesforce.com | 174 | 3,185 | 4,324 |
| paypal.com | 2 | 241 | 1,546 |
| usehatchapp.com | 1 | 136 | 3,266 |
| quartr.com | 4 | 50 | 1,055 |
| dr.dk | 2 | 43 | 2,683 |
| ramp.com | 1 | 18 | 2,170 |
| packagefreeshop.com | 1 | 1 | 542 |
| gldn.com | 1 | 1 | 826 |
| masienda.com | 1 | 1 | 433 |
| principal.com | 26 | 0 | 891 |
| wiley.com | 1 | 0 | 1,932 |
What this does not show, and what to check instead
An empty custom element is not a finding on its own, and the corpus makes the point better than an argument would. All 278 instances of wow-image were empty, and an image wrapper holding no sentence is working correctly. All 40 instances of broadstreet-zone, an advertising slot, were empty on the 7 local news sites carrying it, and so were all 12 instances of lottie-player, which plays an animation. derstandard.at carried 448 instances of which 447 were empty, and served 5,989 words of body text anyway, with only 3 of them inside a component. Emptiness matters when the component is the content and not when it is the furniture, and nothing in a tag name tells you which it is. Running the other way, 15 pages served every word of their body text from inside a custom element, handelsblatt.com among them with 1,612 instances across 101 distinct names and all 2,122 of its words inside them. Components are not the problem. Components holding nothing where the prose should be are.
Four limits apply to every figure above. One page was read per hostname, the home page, so a site that server renders its front page and client renders its article template is counted as clean. No JavaScript was executed, so this is a measurement of what the response contained and never evidence about what any crawler did with it, a distinction that also decided how the iframe measurement had to be written. The 342 hostnames that served no page are absent from every rate, and the 216 answering HTTP 403 are the sites most likely to hold views about crawler access, so the proportions describe the sites that admit an identified crawler. And the corpus is an editorial sampling frame built for platform and industry coverage rather than a random draw, so every number here describes these 1,419 hostnames on one date.
What to do with your own site needs no tool. Fetch the page without a browser, search the response for a lowercase hyphenated tag, and look at what sits between its opening and closing tags. If the answer is nothing, decide whether that component was meant to hold words. If it was, you have a choice between moving the text into the light DOM, adopting the declarative form, or server rendering the route, and the Next.js guidance covers the third for the most common stack. If you want to know which crawlers would even be in a position to run the script, the AI crawler reference lists what each operator publishes, and the guide to being cited by ChatGPT covers what happens after the words are readable. The order matters: readable first, then everything else.
- Fetch without a browser Request the page with no JavaScript engine and keep the raw response. Everything below is read from that one document.
- Find the hyphenated tags Any lowercase tag containing a hyphen is a custom element unless it is one of the eight names the specification reserves.
- Look inside each one Child text is readable by anything. An empty tag means the words were going to arrive later, or not at all.
- Separate furniture from content An empty image wrapper, advertising slot or animation player costs nothing. An empty root element costs the whole page.
- Check for shadowrootmode A template carrying it ships the text in the bytes, but a consumer that drops template subtrees still reads none of it, including this scanner.
Lantad
Published .
A web component is a tag you invented. The HTML parser meets it, does not recognise it, keeps it in the tree as an unknown element, and waits for JavaScript to say what it means. That is the whole design, and it is why the question of whether an AI crawler can read a component page has no single answer: it depends on whether the words were in the response or were going to be assembled later.
Common questions
Do AI crawlers read web components?
It depends on where the component keeps its text, not on the crawler. Text written as child content of the tag is ordinary text that any parser reads. Text that only exists after JavaScript runs is absent from the response, and Google's JavaScript SEO basics page, Last updated 2026-03-04 UTC, states that content not visible in the rendered HTML cannot be indexed. On 8 October 2026, 3,291 of the 7,124 custom element instances Lantad found on 1,077 home pages contained no text in the delivered bytes.
Is an empty custom element a problem?
Only when the component was meant to hold words. On 8 October 2026 all 278 instances of the Wix image wrapper wow-image were empty, which is correct behaviour for an image, and derstandard.at served 5,989 words of body text while 447 of its 448 custom elements were empty. The 8 pages that served zero words with their content behind one empty element are the cases that cost something.
Does declarative shadow DOM fix it?
It puts the component text into the response, which is the part a crawler can act on, and 16 of the 1,077 pages measured on 8 October 2026 used it across 520 templates holding 38,135 words. It is not a complete fix, because any consumer that drops template subtrees reads none of those words. Lantad's own extractor is one such consumer, which is stated in this post rather than left to be discovered.
Does Lantad score custom elements?
No. The scanner has no custom element check, so a page built entirely from components that serve a crawler nothing is caught by the prose parity and word count signals rather than by anything that names the cause. On stanford.edu on 8 October 2026 the extractor read 1,385 words and never opened the 10,436 inside that page's declarative shadow roots, so the gap runs in both directions and is recorded in core/src/renderparity.ts.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.