BlogFindings
GEO optimization: 76 of 225 data-nosnippet attributes sat on an element Google does not support
Google documents four preview controls that decide how much of a page it may use as direct input for AI Overviews and AI Mode, and one of them works on three element types only. Lantad requested robots.txt and then the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 2 October 2026 as LantadBot, following redirects and executing no JavaScript. 1,079 answered HTTP 200 with an HTML content type. 54 of those carried a data-nosnippet attribute, 225 attributes between them, and 76 of the 225 sat on an element type the documentation does not list. On three pages every attribute on the page was one of them.
Google's documentation for AI features in Search, which carries a last updated date of 2025-12-10, names that surface in one sentence: to limit the information shown from your pages in Search, use nosnippet, data-nosnippet, max-snippet, or noindex controls. Three of those four govern snippets rather than indexing, and the robots meta tag specification, last updated 2026-03-24, says nosnippet and max-snippet apply to AI Overviews and AI Mode and limit what may be used as a direct input to them. So this is the documented lever. On 2 October 2026 we asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt and then for their home page, as LantadBot, following redirects and executing no JavaScript, and recorded every one of those four controls on the response. 1,079 answered HTTP 200 with an HTML content type. The result is not that sites set the controls badly. It is that almost nobody sets them, and that a surprising share of the people who did reach for the finest grained one attached it to an element the documentation does not cover.
In short
- GEO optimization has exactly one documented lever over how much of a page Google may feed an AI Overview, and on 2 October 2026 only 2 of 1,079 home pages had pulled it.
- Google's robots meta tag documentation, updated 2026-03-24, supports data-nosnippet on span, div and section elements only. Of 225 attributes Lantad found on 54 home pages on 2 October 2026, 76 sat on something else.
- Two sites account for 64 of those 76: nationalgeographic.com placed 32 on aside elements and linear.app placed 32 on img elements, measured 2 October 2026 and reconfirmed the same day.
- 128 of 1,079 home pages declared max-snippet and 126 of them set it to -1, which Google documents as letting Google choose the length. 645 of the 1,079 carried no robots meta tag and no X-Robots-Tag at all.
| Control | What Google documents it does | Pages | Of 1,079 |
|---|---|---|---|
| nosnippet | No text snippet or video preview, and the content is not used as a direct input for AI Overviews and AI Mode | 1 | 0.09 percent |
| max-snippet | Caps the characters usable as a snippet, and limits how much may be a direct input for AI Overviews and AI Mode | 128 | 11.9 percent, 126 of them set to -1 |
| data-nosnippet | Marks parts of a page as unusable in a snippet, on span, div and section elements | 54 | 5.0 percent, 225 attributes |
| noindex | Removes the page from Search, which removes it from the AI surfaces with it | 16 | 1.5 percent |
What does GEO optimization actually control on your own pages?
Two different mechanisms get collapsed into one word. The first is access, and it is decided before any of this matters: robots.txt says whether a given AI crawler may fetch the URL at all, which is what our robots.txt tester evaluates per token. Of the 1,419 hostnames asked on 2 October 2026, 1,053 returned a parseable robots.txt that allows LantadBot at the site root, 13 disallowed it and were not fetched further, 266 answered a non-200 status, 44 answered with HTML where a text file was expected, and 43 never answered at all.
The second mechanism is what may be shown and reused once the page has been fetched and indexed, and that is where the preview controls live. The distinction is sharp in the documentation and blurred almost everywhere else. A robots.txt rule is a statement about fetching. nosnippet and max-snippet are statements about display and reuse, which is why Google's own text has to spell out that they reach the AI surfaces as well as the blue links. A page can be fully crawlable and still, by its own declaration, be unusable as a quotation. That combination is rare, and we measured how rare.
It also means the two controls fail in different ways. A robots.txt mistake is loud: the page is not fetched, and tools that model AI visibility notice immediately. A preview control mistake is silent. The page is fetched, the page is indexed, the attribute is sitting in the markup, and nothing anywhere reports that it was ignored. Google's AI features page has a troubleshooting section for exactly this, and its first instruction is to make sure the preview control is correct and visible to Googlebot. That is an odd first step to need to publish, and the measurement below suggests why it is there.
Sample Illustrative, not a measurement of any real site.
Flow: Crawler requests the URL to robots.txt allows the path?; robots.txt allows the path? (no) to Never fetched; robots.txt allows the path? (yes) to Page fetched and indexed; Page fetched and indexed to noindex: out of Search; Page fetched and indexed to nosnippet: no direct input; Page fetched and indexed to max-snippet: capped input; Page fetched and indexed to data-nosnippet: part of page; Page fetched and indexed (nothing declared) to Quotable in full.
76 of 225 data-nosnippet attributes sat on an element Google does not support
data-nosnippet is the only one of the four controls that operates on part of a page rather than all of it, which makes it the one a publisher reaches for when the goal is to keep a disclaimer, a paywall notice or a navigation block out of a quotation while leaving the article quotable. The documentation is specific about where it may go: the attribute is supported on span, div and section elements. The same page prints a counter-example in its own code sample, a custom tag carrying the attribute, and labels it not valid because it is not a span, a div or a section.
54 of the 1,079 readable home pages carried the attribute, 225 attributes between them. 149 of the 225 sat on one of the three supported elements: 117 on a div, 29 on a span, 3 on a section. The remaining 76 sat on something else. 33 were on an aside, 32 on an img, 2 each on nav, header and footer, and one each on li, ul, a, fieldset and dialog. By the documentation, those 76 declarations do nothing at all, and nothing in the page or in any report tells their authors so.
Two sites account for 64 of the 76. nationalgeographic.com placed 32 attributes on aside elements and linear.app placed 32 on img elements, both reconfirmed by a second request the same day. On a third page, london.gov.uk, the single attribute on the page sat on a fieldset. Those three are the cases where every data-nosnippet attribute on the home page is on an unsupported element, so the entire attempt is inert; 8 further pages mixed supported and unsupported elements, and 43 used only supported ones. The img count is the most understandable error of the set, because an image is exactly the kind of thing a publisher wants to keep out of a preview, but max-image-preview and noimageindex are the controls documented for that and data-nosnippet is documented as governing textual parts of a page.
The values tell a smaller story in the same direction. 173 of the 225 attributes were written as data-nosnippet="true", 43 were bare, and 9 carried an empty string. All three forms behave identically, because the documentation states the attribute is a boolean attribute and that any value specified is ignored. Writing true is harmless. It does suggest the authors expected a value to be read, which is the same misreading that puts the attribute on an aside.
| Element | Supported by the documentation | Attributes |
|---|---|---|
| div | yes | 117 |
| span | yes | 29 |
| section | yes | 3 |
| aside | no | 33 |
| img | no | 32 |
| nav | no | 2 |
| header | no | 2 |
| footer | no | 2 |
| li, ul, a, fieldset, dialog | no | 5 |
126 of the 128 max-snippet values were the one that changes nothing
max-snippet takes a character count, and the documentation gives two special values: 0 is equivalent to nosnippet, and -1 means Google will choose the snippet length it believes is most effective. A positive number is therefore the only form of the directive that constrains anything, and -1 is a declaration that you accept the default you would have had anyway.
128 of the 1,079 home pages declared max-snippet. 126 of them set it to -1. One set 160 and one set 300. So of 1,079 pages, two had capped the number of characters Google may use as a direct input to an AI Overview, and 126 had written a directive whose effect is identical to writing nothing. 125 of those 126 also sent max-image-preview:large on the same page, and 220 of the 1,079 pages sent that permissive image value at all. A directive that appears 126 times, almost always bracketed by the same companion value, is a template default travelling with a plugin rather than 126 separate decisions about snippet length. It is worth knowing that is what it is, because a site auditing itself will see a snippet directive present and tick the box.
645 of the 1,079 pages carried no robots meta tag and no X-Robots-Tag header at all, so on roughly three fifths of the corpus every one of these controls is simply at its default. This is consistent with a narrower measurement we published on 6 September 2026, when 0 of 81 home pages blocked their own snippet and 3 pages carried max-snippet with all three set to -1. That post asked whether sites were disqualifying themselves from being cited, and answered no. This one asks a different question of a corpus thirteen times the size: among the sites that do reach for the controls, does the declaration do what its author intended. For max-snippet the answer is that it mostly does nothing by design, and for data-nosnippet a third of the attributes do nothing by mistake.
The practical reading for anyone doing answer engine optimization is narrow and worth stating plainly. These controls only ever subtract. Setting max-snippet:-1 does not buy a longer quotation, and no value of it makes a page more likely to be cited in Google AI Overviews. The upside case for touching them is a publisher who genuinely wants less of their text reused; everyone else is editing a field that was already set the way they want it. The pages worth worrying about are the two that capped themselves and the ones where a meta tag reaches a crawler carrying something nobody intended.
The two pages that capped Google, and the header one of them hid it in
Both pages that set a real cap are news publishers, which is the stratum where a licensing position on reuse is most likely to be an actual position. derstandard.at sent a robots meta tag reading max-snippet:160, max-image-preview:large, max-video-preview:-1. france24.com sent no robots meta tag at all and set max-snippet:300, max-image-preview:large, max-video-preview:3 in an X-Robots-Tag response header, which we reconfirmed on two further requests.
That second case is the one worth lifting out, because it is invisible to the obvious audit. Anyone checking a snippet policy by opening the page source will find nothing on france24.com, and will be wrong. 12 of the 1,079 pages sent an X-Robots-Tag header, and that one is the only one of the 12 carrying a positive snippet cap; the other 11 carried noindex, index and follow, noimageindex, the catch-all value all, or max-snippet:-1. A header and a meta tag have equal standing in the documentation, which also notes that Google does not enforce placement in the head and will respect a robots meta tag in the body as well. Checking one location and not the other is how an audit reports a policy a site does not have, and it is the same class of error as reading a page without seeing what the crawler saw. We have written before about an X-Robots-Tag that turned out to be a captcha rather than a decision.
The single nosnippet in the corpus belongs to iitb.ac.in, an education host, and it arrives inside a longer declaration: index, follow, nosnippet, noimageindex, notranslate, max-image-preview:standard. Read against the documentation that is a page asking to be indexed and followed while refusing to supply a snippet, an image preview or a translation, and the nosnippet clause is the part that also withholds the content as a direct input for AI Overviews and AI Mode. No page in the corpus carried both nosnippet and max-snippet, so the documented conflict rule, under which the more restrictive of the two wins, never had to be applied.
One exception deserves to be stated next to these numbers because it limits what a cap achieves. The documentation says the max-snippet limit does not apply where a publisher has separately granted permission for the use of content, and gives in-page structured data as one example of such permission alongside a licence agreement with Google. A news site that caps text at 160 characters and then publishes a full article description in its JSON-LD has granted with one hand what it limited with the other. We did not evaluate whether either of these two sites does that, and the claim here is only about what the documentation says the cap does not cover.
GET as LantadBot, redirects followed, 2 October 2026
- GET https://www.derstandard.at/ 200 text/html
- meta name="robots" content="max-snippet:160, max-image-preview:large, max-video-preview:-1" cap 160
- GET https://www.france24.com/en/ 200 text/html
- no robots meta tag in the delivered HTML none
- X-Robots-Tag: max-snippet:300, max-image-preview:large, max-video-preview:3 cap 300
- remaining 1,077 readable home pages no positive cap
What this measurement does not establish
This is a census of declarations, not of behaviour. We read what 1,079 home pages said about their own snippets on one day. We did not observe Googlebot fetching any of them, we did not see an AI Overview generated from any of them, and we cannot show from this data that Google honoured or ignored a single one of these directives. The claim that nosnippet and max-snippet reach the AI surfaces is Google's claim, published in its own documentation, and our methodology page is explicit about where that line sits between what we measure and what a vendor states.
The sample is home pages, one per host, and a home page is the least likely page on a site to carry an article-level snippet policy. A newsroom that caps snippets on every article template may well leave the front page alone, so the two caps found here are a floor on how many of these 1,419 organisations have a snippet position, not an estimate of it. The 13 hosts that disallowed LantadBot at the root are absent by their own instruction, and a further 327 are absent because 284 answered something that was not HTML with a 200 status and 43 never answered at all. Our research page carries the standing description of the corpus, and what LantadBot requests is published so any of these hosts can check what we asked for.
These four controls are also Google's, and only Google's. Nothing in them is a cross-engine standard: the documentation that defines them is Search documentation, and a site hoping to shape what ChatGPT or Perplexity quotes is not served by any of them. Those engines publish their own crawler tokens and their own robots.txt expectations, which is a question of access rather than of preview length, and we have measured separately that noindex is named by only 2 of 9 operator documentation pages. The general lesson that a control has to exist on the surface you care about is the same one behind the AI Overviews opt out not living on your site.
Finally, a measurement like this says nothing about whether any of it is the right thing to spend attention on. For most sites the honest answer from this data is that the preview controls are already set the way they want them, and that the time is better spent on whether the page reaches a crawler as prose at all, which is where we keep finding the large failures: 121,312 of 1,269,054 words sat behind a hidden marker in a comparable corpus read in September, and a canonical tag pointed somewhere the page was not served on 63 of 1,069 pages. A framework that renders its text server side, which our Next.js guide covers, moves more of this than any snippet directive will.
-
Declarations on 1,079 home pagesEstablished Counted from the delivered bytes and response headers: 128 max-snippet, 1 nosnippet, 54 pages with data-nosnippet, 16 noindex. -
76 of 225 attributes on unsupported elementsEstablished Measured against the element list in Google's documentation, and the three all-inert pages were reconfirmed by a second request. -
Whether Google obeyed any of theseNot measured No crawl by Googlebot was observed and no AI Overview was generated. The effect of each control is Google's published claim. -
Article-level snippet policyOut of scope One home page per host. A cap applied on article templates only would not appear in this count. -
Any engine other than GoogleNot applicable These four controls are defined only in Google Search documentation and no other engine documents honouring them.
Lantad
Published .
A site that wants to be read by an answer engine spends most of its effort on things no specification governs: how the prose is written, what the page is about, whether an entity is named consistently. Those matter, and none of them is a setting. There is a smaller surface that is a setting, written down, and enforced by the crawler rather than inferred by a model, and GEO work that skips it is arguing about taste while leaving a switch unflipped.
Common questions
Does data-nosnippet work on any HTML element?
No. Google's robots meta tag documentation, last updated 2026-03-24, supports the attribute on span, div and section elements, and its own code sample marks a custom tag carrying the attribute as not valid. Of 225 attributes Lantad found across 54 home pages on 2 October 2026, 76 sat on another element type, most often an aside or an img.
What does max-snippet:-1 do?
Nothing that not setting it would not also do. The documentation defines -1 as Google choosing the snippet length it believes is most effective, which is the default behaviour. 126 of the 128 home pages declaring max-snippet on 2 October 2026 used that value, and 125 of the 126 paired it with max-image-preview:large.
Do the snippet controls affect AI Overviews?
Google says so in its own documentation. The robots meta tag page states that nosnippet applies to AI Overviews and AI Mode and prevents the content being used as a direct input for them, and that max-snippet limits how much may be used as a direct input. Lantad has not observed that behaviour and reports it as Google's published claim rather than a measurement.
Can a snippet cap be set without appearing in the page source?
Yes. An X-Robots-Tag response header has the same standing as a robots meta tag. On 2 October 2026 france24.com sent max-snippet:300 in that header and no robots meta tag in its HTML, so an audit reading only the page source would have found no policy at all.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.