BlogFindings
Semantic HTML for AI: 379 of 1,066 home pages declared no main landmark
Lantad asked all 1,419 hostnames in this repository's committed corpus for robots.txt on 9 October 2026 and read each home page it was allowed to read, with no JavaScript executed. 1,066 answered HTTP 200 with HTML. 687 declared at least one main element, 379 declared none, and 301 declared neither a main nor an article. Across the 1,066 pages the shipped extractor classified 1,335,841 words, of which 418,642 were boilerplate inside nav, header, footer or aside.
That is what semantic HTML for AI comes down to. Not a schema, not a file in the root, not a header: the question of whether the markup names its own content. The HTML Standard provides exactly one element for the job, and both of the specifications that define it say the same thing in almost the same words. So this run went and counted it, across every hostname in this repository's committed corpus, using the extraction ruleset the scanner ships rather than a script written for the occasion.
687 of 1,066 readable home pages declared a main element. 379 declared none, 301 declared neither a main nor an article, and the furniture elements were better adopted than the content element on every count: 857 pages declared a footer and 853 declared a nav. The pages were more willing to name what does not matter than to name what does.
In short
- Semantic HTML for AI starts with one element, and of the 1,066 corpus home pages Lantad read on 9 October 2026, 687 declared a main landmark and 379 declared none.
- 301 of the 1,066 home pages measured on 9 October 2026 declared neither a main nor an article element, which leaves an extractor nothing to separate content from furniture except the paragraph, the list item and the heading.
- Of the 1,335,841 words classified across those 1,066 pages, 418,642 sat inside nav, header, footer or aside, and on 237 pages that furniture outweighed the content.
- The page builder decided the landmark more often than the author did: 53 of 54 Wix and Squarespace home pages declared a main element on 9 October 2026, against 7 of 42 built on Bubble and comparable no-code tools.
- Of the 686 pages carrying a main element that were read a second time, 20 carried more than one non-hidden main and 27 placed a main under an ancestor the HTML Standard does not permit.
What does semantic HTML for AI actually decide?
It decides the denominator. Everything downstream of extraction, from prose parity to a word count to whether a claim appears in the text at all, is computed over whatever the extractor kept. Get that boundary wrong and every number after it is wrong in the same direction, which is why the ruleset is in one file and is the same file in development, in the command line tool and in production.
The ruleset is short enough to state in full. Script, style, noscript, template, svg and iframe subtrees are dropped before anything is counted. Anything inside a main or an article element is content. Anything inside a nav, header, footer or aside is furniture, and furniture wins a conflict, so a paragraph inside a footer stays furniture. Outside any of those, a paragraph, a list item or a heading is still content, and everything else is neutral: present on the page, not counted as either. Those four lines are the whole instrument, and the measurement below is what happens when they meet the web. Note what the ruleset is not doing: naming the content of a document is a different job from naming the entities in it, which is what structured data is for, and a page can do either one without the other.
The two elements that carry the most weight in that list are the two the specifications define most narrowly. MDN's reference for the element says the main element represents the dominant content of the body of a document, and that a document must not have more than one main element without the hidden attribute. The W3C's WAI-ARIA 1.2 Recommendation, published 6 June 2023, defines the matching role as a landmark containing the main content of a document and states that an author should mark no more than one element with it. The normative source for the element itself is the WHATWG HTML Standard, whose text is at html.spec.whatwg.org/multipage/grouping-content.html#the-main-element and whose host this site does not link; it requires that a document must not have more than one main element that does not have the hidden attribute specified, and that every main element be hierarchically correct, meaning its ancestors are limited to html, body, div, a form without an accessible name, and autonomous custom elements.
Two specifications, one rule, stated twice. The last section of this post counts how often it held.
Flow: A block of text to In script, style, svg, iframe?; In script, style, svg, iframe? (yes) to Dropped; In script, style, svg, iframe? (no) to In nav, header, footer, aside?; In nav, header, footer, aside? (yes) to Furniture; In nav, header, footer, aside? (no) to In main or article?; In main or article? (yes) to Content; In main or article? (no) to In a p, li or heading?; In a p, li or heading? (yes) to Content; In a p, li or heading? (no) to Neutral.
How many home pages declared a main landmark?
687 of 1,066, which is 64.4 percent, and the shape of the remainder matters more than the headline. 342 pages declared at least one article element, which the extractor treats as a content container for the same reason: it is defined as a self contained composition intended to be independently distributable. 78 of those 342 declared an article without a main, so counting both landmarks together rescues 78 pages and leaves 301 that named their content with neither.
Set that against the furniture. 857 of the 1,066 pages declared a footer, 853 declared a nav, 783 declared a header and 758 declared a section. Only 130 declared an aside. A page is markedly more likely to tell a machine where its navigation is than where its content is, and the gap is not small: 857 against 687 on the same 1,066 pages.
There is an explanation for that ordering which is worth stating because it is probably correct and is not an excuse. Nav, header and footer are structural containers an author reaches for while building a layout, and a layout has a header whether or not anyone is thinking about machines. Main is the element you add when you have thought about the document as a document. It is the one in the list that requires an intention, which is exactly why it is the one worth counting, and why its absence is informative rather than merely untidy.
Nothing here says a page without a main element is unreadable. The extractor still kept a median of 349 words from those 379 pages, because the paragraph, list item and heading rules carried them. It says the page handed the decision to a heuristic, and the fifth section measures what that costs when the heuristic has nothing to work with. We have measured neighbouring versions of this problem before, including 311 of 704 article pages that ran over 300 words with no heading at all, and the two failures compound: a page with no landmark and no headings has only its paragraphs left.
The page builder decided the landmark, not the author
The platform strata settle the question of where the element comes from, and the answer is that on most sites nobody chose it. 53 of 54 Wix and Squarespace home pages declared a main element. 33 of 34 Shopify storefronts did, which is a theme decision rather than a merchant one and is edited in the same place on a Shopify theme as on any other template. 7 of 42 sites built on Bubble and comparable no-code builders did. Those figures are not three populations of authors behaving differently. They are decisions taken once inside each platform's own templates and then inherited by every site built on it.
Read down the platform table and the ordering tracks how much of the document each tool authors on the publisher's behalf. The hosted website builders that emit a whole page from a template score highest. The static documentation generators follow at 26 of 31. The tools that hand the author a blank canvas and absolute positioning score lowest: 12 of 30 on Framer, 15 of 42 on Webflow, 7 of 42 on the no-code builders. A canvas has no document outline to inherit, so unless the builder inserts the landmark, nothing does.
The industry strata say the same thing from the other side, with a narrower spread because industry does not determine a stack. 93 of 116 software-as-a-service home pages declared a main element, the highest of the eight, and 32 of 61 ecommerce home pages did, the lowest. Commerce is where the hosted platforms concentrate, so the low ecommerce figure is not a contradiction of the Shopify figure: the ecommerce stratum is sampled by sector rather than by platform and contains a good deal of bespoke retail.
For anyone acting on this, the useful consequence is that the fix is almost never a content project. On a templated site it is one edit in one layout file, and it is the same edit whether the stack is React, Next.js or a generated app. We have measured what the same class of platform decision does elsewhere: 10 of 42 no-code home pages served a crawler zero words, and the landmark figures put that result in a wider frame.
| Platform stratum | Declared a main | Pages read | Share |
|---|---|---|---|
| Wix and Squarespace | 53 | 54 | 98 percent |
| Shopify direct to consumer | 33 | 34 | 97 percent |
| Static documentation sites | 26 | 31 | 84 percent |
| Local and regional media | 21 | 29 | 72 percent |
| SaaS marketing sites | 25 | 35 | 71 percent |
| Single page app startups | 25 | 44 | 57 percent |
| Framer | 12 | 30 | 40 percent |
| WordPress small business | 13 | 35 | 37 percent |
| Webflow | 15 | 42 | 36 percent |
| Bubble and other no-code | 7 | 42 | 17 percent |
Where the words went: 418,642 of 1,335,841 were furniture
Running the shipped ruleset over all 1,066 pages classified 1,335,841 words. 803,062 of them were content, 418,642 were furniture inside a nav, header, footer or aside, and 114,137 were neutral: text in a div or a span outside any landmark and outside any paragraph, list item or heading. Just under a third of every word classified across the corpus exists to help a human navigate, and an extractor that counted those words would be reporting a site's own menu back to it as its message.
That ratio is not a new claim, and the corpus is what is new about it. We first counted it on five captured pages, where 1,395 of 2,729 text blocks were navigation rather than main content while holding under a fifth of the text, and the question left open was whether five pages were representative. On 1,066 they are, at least in direction.
The medians are lower than those totals suggest and are the more useful figures: 564 content words, 185 furniture words and 7 neutral words per page. The distribution is long tailed at both ends. On 237 of the 1,066 pages the furniture outweighed the content outright. On 68 pages the extractor kept no content words at all.
The extreme cases are worth naming because they show what the ratio looks like when it fails completely rather than gradually. derstandard.at served 10 content words against 6,224 words of furniture, and a further request the same day showed the mechanism: 6,159 of the 6,168 words sitting inside its main element were also inside one of the 257 header, footer and nav elements nested within that landmark, and furniture wins the conflict. allbirds.com served 260 against 14,392. canadiantire.ca served 125 against 10,657 and declared no main element at all, and oliverbonas.com served 16 against 1,071 on the same pattern. Three of those four are retail, which is consistent with the ecommerce stratum result above and with the fact that a storefront home page genuinely is mostly navigation. The question that leaves is whether the page says anything at all about the business, and on these four the measured answer is close to no.
This is the figure a reader can check on their own site without any tooling from us. Count the words inside your nav, header and footer elements, count the words outside them, and compare. If the first number is larger, that is what an extractor sees, and a scan at what GPTBot sees will say the same thing with the request log attached. The related failure, where the words are present but hidden from a non-rendering client, we have measured separately: 121,312 of 1,269,054 words sat behind a marker on an earlier corpus pass.
| Class | Words | Share | Median per page | What the ruleset means by it |
|---|---|---|---|---|
| Content | 803,062 | 60.1 percent | 564 | inside main or article, or in a p, li or heading anywhere |
| Furniture | 418,642 | 31.3 percent | 185 | inside nav, header, footer or aside, which wins any conflict |
| Neutral | 114,137 | 8.5 percent | 7 | visible, in none of the above, counted as neither |
Without a landmark, only paragraphs, list items and headings survive
The 301 pages that declared neither a main nor an article are the cohort where the ruleset has run out of structural instructions. Everything the extractor kept from them, 156,940 words in total, had to arrive through the one rule that does not need a landmark: that a paragraph, a list item or a heading is content wherever it sits. There is nothing else left to consult.
That rule carries a lot. Across all 1,066 pages, 424,525 of the 803,062 content words were anchored by a paragraph, 75,868 by a list item and 139,227 by a heading at some level, with h3 alone accounting for 67,142. A further 163,442 words sat inside a landmark without being inside any of those elements, which is precisely the text that only the landmark rescued, and on a page with no landmark that text is not content at all.
So the cohorts separate. The 687 pages with a main element yielded a median of 625 content words. The 379 without one yielded 349. The 301 with neither landmark yielded 239. And at the floor, 65 of those 301 pages yielded zero content words, against 3 of the 687 pages that declared a main element.
Read that last pair carefully, because the obvious reading of it is wrong. Declaring a main element does not create prose. A page with no landmark and no prose is usually a page that renders its text with a client-side framework, or lays it out in positioned divs, and the missing landmark is a symptom of the same build rather than the cause of the empty result. What the comparison does establish is the size of the margin the landmark provides when the page is otherwise borderline, and the direction is not in doubt: the element is the difference between an extractor reading a document and an extractor guessing at one. The only direct test of that margin we have run is the one where we deleted the element ourselves and held everything else constant, and removing main cost 1,665 of 13,615 words across five captured pages, worth 43 percent of the corpus on one of them.
Three of the 1,066 pages managed both at once, declaring a main element and still yielding no content words: hku.hk, skyscanner.net and thegrilledcheesefactory.fr. A further request the same day separated two different reasons. On skyscanner.net the landmark held no text at all in the bytes the server sent. On hku.hk every one of the 348 words inside the landmark was also inside one of the three boilerplate elements nested in it, and on thegrilledcheesefactory.fr all six words sat inside a nested nav. This run rendered nothing, so what arrives in those elements once a browser runs the page is untested here, and the canary pages exist to demonstrate that general case. A landmark whose entire contents are furniture reports the same figure as an empty one.
687 pages with a main element
- Median content words: 625
- Median furniture words: 235
- Furniture outweighed content: 153 pages
- Kept zero content words: 3 pages
- The landmark alone rescued 163,442 words corpus wide
301 pages with neither landmark
- Median content words: 239
- Median furniture words: 46
- Furniture outweighed content: 67 pages
- Kept zero content words: 65 pages
- Every kept word had to be in a p, an li or a heading
30 of 686 pages broke a rule the HTML Standard sets
Adoption is one question and correctness is another, so the 687 pages carrying a main element were read a second time and parsed for the structure of the element rather than its presence. 686 answered again and parsed. 656 of those were clean on both requirements: exactly one non-hidden main element, and every main element sitting under permitted ancestors only. 30 broke at least one.
20 pages carried more than one non-hidden main element, which is the requirement both specifications state in the same breath. Not one page in the corpus used the hidden attribute on a main element, so there is no page where a second main is the legitimate hidden case the standard carves out. hookagency.com carried ten. huggingface.co carried six, acc.org five, and foxnews.com, siigo.com and hiutdenim.co.uk four each. A document with ten main elements has not named its content; it has applied a class name that happens to be an element.
27 pages placed at least one main element under an ancestor the standard does not permit. Counting the pages by the offending ancestor: 23 under a section, 22 under another main, 11 under an article, 5 under an anchor, 3 under an aside, 2 each under a header, a footer and a span, and 1 inside a template. The ancestor list in the specification is deliberately tiny, and the reason is the one this post is about: a landmark nested inside other content is not a landmark, because it no longer identifies the document's dominant content. 10 of the 27 had a main inside another main, which is the same defect as the multi-main count seen from the inside.
None of this is catastrophic on its own, and that is the point worth making rather than inflating. The extractor handles all of it: nested landmarks just mean the content flag is already set, and a main inside a footer stays furniture because furniture wins. A reader should take these 30 pages as evidence about where the markup came from rather than as a list of broken sites. An element emitted ten times on one page was emitted by a tool, and a tool that gets the count wrong is a tool that is not reading the specification, which is a reasonable thing to know about the software generating your pages.
| Finding | Pages | What the specification requires |
|---|---|---|
| Exactly one non-hidden main, all ancestors permitted | 656 | the compliant case |
| More than one non-hidden main element | 20 | a document must not have more than one without the hidden attribute |
| A main under a disallowed ancestor | 27 | ancestors limited to html, body, div, an unnamed form, custom elements |
| A main nested inside another main | 10 | a subset of both rows above |
| A main carrying the hidden attribute | 0 | the one case where a second main is permitted |
What this measured, and what it did not
One page per hostname, and that page the home page. A home page is the least representative page on most sites for this particular question, because it is the page most likely to be mostly navigation by design. The landmark adoption figures should travel to interior pages, since they are a property of the template, but the furniture share almost certainly should not, and this run did not read an interior page to find out.
No JavaScript was executed. Everything above is the markup as served, which is the view a non-rendering client has, and the three empty-landmark pages are the clearest demonstration of how far that view can diverge from a browser's. A page that inserts its main element from a script was counted as having none. That is the correct measurement of what a non-rendering AI crawler receives and it is not a measurement of what the page is.
The second pass parsed the served markup directly rather than constructing a DOM the way a browser would. A browser repairs unclosed elements and implied closures before anything reads the tree, so a handful of the 27 ancestor findings may be artefacts of malformed markup rather than of genuine nesting. The multi-main count does not depend on that, because it only requires counting opening tags.
No vendor documents what any of this does. No crawler operator we have read publishes that it uses the main element, which means the honest claim is narrow and worth stating exactly: these figures describe what Lantad's extractor does with the corpus, and the specifications describe what the element means. Whether GPTBot, ClaudeBot or an AI Overview weights the landmark is not something this run tested, and treating an extractor's behaviour as a search engine's would be the same error as treating a scoring weight as a finding. The methodology page states the general rule and the bot page documents the agent that made the requests.
Finally, the denominator is the pages that answered, not the pages that exist. 353 of the 1,419 hostnames produced no readable home page: 215 answered 403, 44 never answered the robots.txt request at all, 39 answered it with a 5xx, and 13 disallowed LantadBot at the root. Sites behind strict bot defences are over-represented in that 353, and they are not a random sample of the web. The same selection effect runs through every corpus pass this site publishes and is set out at length in the crawlability study. Everything above is measured on the 1,066 that let a declared crawler read a page, which is a different population from the web and the only one an AI visibility measurement ever gets.
-
Landmark adoptionMeasured 687 of 1,066 home pages declared a main element, 301 declared neither main nor article. -
Word classificationMeasured 418,642 of 1,335,841 classified words were furniture, and 237 pages carried more furniture than content. -
Specification complianceMeasured 30 of 686 pages broke a main element requirement, 20 of them by declaring more than one. -
Interior pagesNot measured One page per hostname. The furniture share in particular should not be read across to interior pages. -
Rendered markupNot measured No JavaScript executed, so a landmark inserted by a script was counted as absent. -
Effect on any engineNot measured No crawler vendor documents using the element, and no engine was asked anything in this run.
Lantad
Published .
An extractor has one hard problem, and it is not reading. It is deciding which of the words on a page are the page. A home page arrives carrying a cookie notice, a mega menu, a language switcher, a cookie notice again inside a modal, four columns of footer links and, somewhere in the middle, the two paragraphs that say what the company does. Something has to draw the line, and the only instructions the document offers are the element names the author chose.
Common questions
What is semantic HTML for AI?
It is the practice of choosing element names that tell a machine what each part of a page is, rather than styling anonymous containers to look right. For an AI crawler the decisive case is the boundary between content and navigation, because an extractor has to draw that line before it can count anything. The HTML Standard provides one element for the content side of it, main, and 687 of 1,066 corpus home pages declared it on 9 October 2026.
Does a missing main element mean an AI crawler reads nothing?
No. Lantad's extractor still kept a median of 349 words from the 379 home pages that declared no main element, because a paragraph, a list item or a heading counts as content wherever it sits. The cost shows up in the margin rather than the headline: the 301 pages declaring neither main nor article yielded a median of 239 content words against 625 on pages that declared a main, and 65 of those 301 yielded none at all.
How many main elements may a page have?
One that is not hidden. The W3C's WAI-ARIA 1.2 Recommendation of 6 June 2023 says an author should mark no more than one element with the main role, and the HTML Standard requires that a document not have more than one main element without the hidden attribute. 20 of the 686 corpus pages measured on 9 October 2026 carried more than one non-hidden main, and none of the corpus pages used the hidden attribute on a main element at all.
How do I check my own page without any tooling?
View the source your server sends rather than the inspector's tree, search it for a main element, and if there is one check that nothing but html, body, div or a form sits above it. Then count the words inside your nav, header and footer elements against the words outside them. On 237 of the 1,066 corpus home pages read on 9 October 2026 the first number was larger, and that ratio is what an extractor works from.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.