BlogFindings

Google AI Mode: 132 of 1,083 home pages sent the directive that governs it, and 3 limited anything

Google's robots meta tag documentation states that nosnippet prevents content being used as a direct input for AI Overviews and AI Mode, and that max-snippet limits how much may be used. Lantad read the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 7 October 2026 with no JavaScript executed. 1,083 answered HTTP 200 with HTML. 132 sent one of those two directives. 129 of them set max-snippet to -1, which Google defines as no limit at all.

16 min read Lantad

Google's robots meta tag documentation, carrying Last updated 2026-03-24 UTC when it was opened for this post, attaches the consequence in two places. Of nosnippet it says the rule will also prevent the content from being used as a direct input for AI Overviews and AI Mode. Of max-snippet it says the rule will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode. That is a page-level lever over a generative answer surface, documented by the vendor, readable by anybody, and requiring no account. So this run counted who uses it. Lantad read the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 7 October 2026 as LantadBot, following redirects, with no JavaScript executed. 1,083 answered HTTP 200 with an HTML content type. 132 of those sent a snippet directive, and three sent one that limits anything.

In short

  • Google AI Mode has one documented on-page control, and it is the snippet directive: Google's robots meta tag documentation, Last updated 2026-03-24 UTC, states that nosnippet will also prevent the content from being used as a direct input for AI Overviews and AI Mode, and that max-snippet will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode.
  • Lantad measured 132 of 1,083 readable home pages sending a snippet directive on 7 October 2026, and exactly 3 of those 1,083 sent a value that limits anything: iitb.ac.in sent nosnippet, derstandard.at sent max-snippet:160, and france24.com sent max-snippet:300 in an X-Robots-Tag header.
  • 129 of the 132 pages set max-snippet to -1, which Google's documentation defines as Google choosing the snippet length it believes is most effective. It is the default behaviour written out, so it restricts nothing. No page in the corpus sent max-snippet:0.
  • The directive looks generated rather than chosen: 122 of its 136 occurrences carried the identical five-token set of index, follow, max-image-preview:large, max-snippet:-1 and max-video-preview:-1, differing only in order and spacing, and 89 of the 129 pages carried a marker for Yoast, Rank Math, All in One SEO or SEOPress.
  • Google-Extended is not the AI Mode lever, though it is the token most often reached for: Google's crawler documentation, Last updated 2026-07-14 UTC, scopes that token to training and grounding in Gemini Apps and Vertex AI, and names neither AI Overviews nor AI Mode.
StageHostsWhat happened
Hostnames asked1,419392 in ten platform strata, 1,027 in eight industry strata
Never returned a status3423 failed in the client, 10 hit the timeout, one terminated
Answered HTTP 403217Refused this crawler
Answered HTTP 50357Served no page to this client
Answered some other status289 of 202, 7 of 429, 5 of 404, 2 of 406, one each of 400, 401, 451, 498, 500
Answered 200 with HTML1,083The denominator for every rate below
Sent no crawler directive at all651No robots meta, no vendor meta, no X-Robots-Tag header
Sent at least one crawler directive432420 in a meta element only, 7 in a header only, 5 in both
Sent a snippet directive132nosnippet or max-snippet
Set max-snippet to -1, meaning no limit129Restricts nothing
Sent a value that limits anything3iitb.ac.in, derstandard.at, france24.com
Sent max-snippet:00Nobody
One GET of https://<host>/ per hostname as LantadBot/1.0 (+https://lantad.co/bot), redirects followed, 20 second timeout, no JavaScript executed, from one network location. Measured by Lantad on 7 October 2026 across the 1,419 hostnames in this repository's two committed corpus seed files.

What controls whether your content appears in Google AI Mode?

Two things, read at different moments, and only one of them is on your page. Google's crawler documentation is explicit about the first. Google's page on its common crawlers, Last updated 2026-07-14 UTC, states that Google-Extended is a standalone product token publishers can use to manage whether content Google crawls may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini, and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. The same entry states that Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. AI Mode is a Google Search surface. AI Overviews is a Google Search surface. Neither is named in that scope.

The second is the snippet directive, and it is the one the robots meta tag page ties to both surfaces by name. This matters because the two controls are not substitutes and are not even about the same thing. Google-Extended governs Gemini's training and the grounding of Gemini Apps. The snippet directives govern how much of your text Google may take as a direct input when it composes an answer in Search. A site can block Google-Extended at the root and still be quoted at length inside an AI Overview, which is the configuration this blog described when it reported that AI Overviews retrieved less from Google-Extended blockers without their disappearing from the surface, and the distinction is the same one behind how to get cited by Gemini through one token covering two systems.

There is a third control and it is not on your site at all. Google began testing a Search Console setting in June 2026 that excludes a property from AI Overviews and AI Mode, which this blog wrote about under the finding that the AI Overviews opt out does not live on your site. It is an account setting inside an authenticated dashboard, so no external scan can see whether it is set, including this one. Everything measured below is therefore the on-page layer only, and a site could be opted out at the account level while sending max-snippet:-1 on every page. The two statements do not contradict each other and this measurement cannot see the first.

What a reader can take from the documentation, before any count, is that the on-page lever is narrow and precise. It does not say a crawler will stay away. It says the text will not be used as a direct input, or will be used only up to a character budget you set. That is a different kind of control from the gate that two layers decide if AI can read your site describes, and it sits strictly after it.

ControlWhere it livesScope its documentation statesCovers AI Mode
nosnippet and max-snippetMeta robots tag or X-Robots-Tag headerSnippet length, and use as a direct input for AI Overviews and AI ModeYes, named
Google-Extendedrobots.txt user-agent tokenGemini model training and grounding in Gemini Apps and Vertex AINot named
Search generative AI settingSearch Console account, not the siteExcludes a property from AI Overviews and AI ModeYes, per Google's announcement
Three controls over Google's generative surfaces, with the scope each one's own documentation gives it, read at source on 7 October 2026. Reported from Google's published statements, not measured by Lantad: no crawler was observed honouring or ignoring any of these.

How many sites set a snippet limit at all?

132 of the 1,083 readable home pages sent either nosnippet or max-snippet. That is the headline count and it overstates the position badly, because 129 of those 132 set max-snippet to -1. Google's documentation defines -1 as a special value meaning Google will choose the snippet length that it believes is most effective to help users discover your content and direct users to your site. It is the default behaviour spelled out as an instruction. A page sending it has told Google exactly what Google was going to do anyway, which means the count of pages in this corpus that have used the documented AI Mode lever to reduce anything is three.

The three are worth naming individually, because each used a different mechanism. iitb.ac.in, the Indian Institute of Technology Bombay, sent index, follow, nosnippet, noimageindex, notranslate, max-image-preview:standard in a single meta named robots. It is the only nosnippet in the whole corpus, and under Google's documented behaviour it is the only home page here whose text is withheld from AI Overviews and AI Mode as a direct input, while still asking to be indexed and followed. derstandard.at sent max-snippet:160 in a meta robots tag, on a home page carrying 6,047 words of extractable text. france24.com sent max-snippet:300, not in a meta tag at all but in an X-Robots-Tag response header, on a page carrying 3,077 words.

Nobody sent max-snippet:0. Google's page lists 0 as a special value meaning no snippet is to be shown, equivalent to nosnippet, so there were two documented routes to the same outcome and 1,083 pages took the first route once and the second route never. The same page notes that where rules conflict the more restrictive applies, giving the example that a page carrying both max-snippet:50 and nosnippet has the nosnippet applied. No page in this corpus created that conflict, because no page sent two different snippet values.

The distribution across the sampling strata is the first clue about why the figure looks the way it does. The wordpress-smb stratum sent a snippet directive on 31 of its 37 readable pages, the highest rate anywhere in the corpus by a wide margin. media-local sent one on 16 of 30 and news on 15 of 62. At the other end, the framer, webflow and shopify-dtc strata sent a snippet directive on none of their 31, 42 and 34 readable pages respectively. That is not a story about editorial policy on AI. It is a story about which content management system writes the tag for you, and it is the same shape this blog found when counting what meta tags for AI search 1,083 home pages actually sent.

  • max-snippet:-1 129 hosts Google chooses the length, so nothing is limited
  • max-snippet:300 1 hosts france24.com, in an X-Robots-Tag header
  • max-snippet:160 1 hosts derstandard.at, on a 6,047 word page
  • nosnippet 1 hosts iitb.ac.in, the only one in the corpus
  • max-snippet:0 0 The documented equivalent of nosnippet, unused
  • Any snippet directive 132 hosts 12.2 percent of the 1,083
  • Any crawler directive 432 hosts The wider population the 132 sit inside
  • No directive of any kind 651 hosts Google's defaults apply untouched
Snippet directive values across the 1,083 readable home pages, counted by host. Measured by Lantad on 7 October 2026. The -1 row is the value Google's documentation defines as no limit.

Why 129 sites set max-snippet to -1

Because something else wrote it. This is an inference rather than a measurement of intent, and the evidence for it is strong enough to state with its limits attached. The 136 occurrences of max-snippet in the corpus are not 136 independently chosen values. 122 of them carry the identical set of five directives, index, follow, max-image-preview:large, max-snippet:-1 and max-video-preview:-1, differing only in the order of the tokens and the presence of spaces after the commas. 79 occurrences are the same string byte for byte: index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1.

To test where that string comes from, every one of the 129 hosts sending max-snippet:-1 was fetched a second time and its HTML searched for platform markers. 98 of the 129 carried WordPress markers. 89 carried a marker for one of four WordPress SEO plugins: 73 for Yoast, 15 for Rank Math, one for All in One SEO and one for SEOPress. 67 of those 89 sent the byte-identical canonical string. The marker proves the plugin is present on the page, not that the plugin wrote the directive, and that is the honest limit of the claim. But a directive that appears as the same five tokens in the same order on 79 pages, overwhelmingly on one platform, is generated output rather than a decision taken 79 times.

This is the part that matters for anybody reading their own markup. A max-snippet:-1 in your head is not evidence that somebody considered AI Mode and chose to allow unlimited extraction. It is considerably more likely to be the default your generative engine optimization stack shipped, in the same way that the 233 of 241 max-image-preview values this blog counted as set to large in its 5 October 2026 census were a default rather than an act. The directive is doing nothing, which is fine if nothing is what you want. The failure mode is believing it is doing something.

Four pages sent max-snippet:-1 to the wrong audience, which is a smaller finding but a clean illustration of markup copied rather than written. highlandparkveterinarian.com, cri.dev, metalbear.com and amref.org each sent it inside a meta named bingbot. Google reads a meta named robots, googlebot or googlebot-news, so a directive addressed to bingbot is not a Google instruction and was excluded from the counts above. It is the same texture of inherited, unexamined markup that this blog documented when 232 of 1,059 robots.txt files carried a defect, and it is why a count of pages sending a directive is never the same as a count of pages exercising a control.

Marker found in the HTMLPages of 129What it indicates
WordPress98wp-content, wp-includes or wp-json referenced
Any of the four SEO plugins below8969 percent of the 129
Yoast SEO73The largest single group
Rank Math15Second largest
All in One SEO1One page
SEOPress1One page
Carried none of those markers27Hand written or written by something else
Sent the byte-identical canonical string79Occurrences, not hosts
Platform and plugin markers in the HTML of the 129 home pages sending max-snippet:-1, each page re-fetched on 7 October 2026. A marker shows the software is present on the page; it is not proof that the software wrote the directive.

Where a snippet directive never gets read at all

Google's robots meta tag page states the precondition before it defines a single directive: these settings can be read and followed only if crawlers are allowed to access the pages that include these settings. A snippet directive is an instruction inside a response body. A crawler that is refused the response never holds the instruction. The two surfaces compose in one direction, the gate first and the page second, and the group selection rules that decide the gate are specified in RFC 9309.

That ordering is not theoretical in this corpus. 217 of the 1,419 hostnames answered HTTP 403 to a plain crawler and 57 answered 503, so 274 hosts served no page at all to this client and whatever directives sit in their heads were unreadable from here. A further 34 never returned a status. None of those 308 hosts appears in any rate above, and a host that refuses an identified crawler outright is exactly the kind of host most likely to have opinions about snippet extraction, so the surviving 1,083 lean toward sites that admit crawlers. The direction of that bias is worth stating: it probably undercounts restrictive configurations rather than overcounting them.

The timing matters too, in the other direction. A directive added to a page today is read the next time the page is fetched, which is a different question from when a robots.txt edit reaches a crawler, and both are slower than a dashboard toggle. There is also a text-level version of the same control that this measurement treats separately. 52 of the 1,083 pages carried a data-nosnippet attribute somewhere in the body, which marks individual spans rather than the page, and 12 of the 132 pages sending a page-level directive also carried one. The attribute has its own failure mode, which this blog measured when it found 76 of 225 data-nosnippet attributes sitting on an element Google does not support, since Google's documentation restricts it to span, div and section elements.

One more asymmetry deserves naming. Of the 432 pages sending any crawler directive, the overwhelming majority ask for more presence rather than less, which is the same conclusion the noarchive census reached from a different angle five days ago. Only a handful of directives in this corpus withhold anything, and the ones that do are often doing housekeeping rather than making a statement: this blog found that five of 391 home pages sending X-Robots-Tag noindex were all a captcha, and the lesson transfers. Counting the presence of a directive is easy. Reading intent from it is not available, and this post does not try. You can check what your own site serves with the robots.txt tester, see what an AI crawler is actually handed with what GPTBot sees, and read the token reference on the AI crawlers page.

Why a snippet directive sits behind the gate. Drawn from the precondition Google's robots meta tag documentation states and the group selection rules in RFC 9309; this is an explanation of the protocol rather than a measurement of any crawler.

What this measurement does not show

No crawler was observed. Nothing here is evidence that Google honoured a nosnippet, that any page was or was not used as a direct input to an AI Mode answer, or that the three sites limiting their snippets achieved anything by it. The statements about what the directives do are Google's published claims about its own products, read at source on 7 October 2026 and reported as claims, which is the only standing they have in this post. Lantad does not score these directives either: the composite grade is built from parity, access, structure and schema, so a clean report from this scanner says nothing about whether a page carries nosnippet, exactly the gap the Microsoft AI opt out post named and which has not closed.

The scope limits are specific. One page was read per hostname, the home page. A publisher that sets max-snippet on its articles and leaves the front page open is invisible to this run, and for a news organisation that is the more likely configuration, so the news and media strata are probably undercounted here. No JavaScript was executed, so a directive injected client side is absent, though a snippet directive added after load is of doubtful value given that the crawler reads the served HTML. The word counts quoted for derstandard.at and france24.com are extractable words from the served HTML with script and style content removed, not a character count of what Google would select, so they are context for the size of the cap rather than a measure of what the cap withholds.

The corpus is also not a sample of the web. It is an editorial sampling frame assembled for platform and industry coverage, with 392 hostnames in ten platform strata and 1,027 in eight industry strata, so every rate above describes these 1,419 hostnames on one date and nothing wider. The strata are listed on the crawlability study page for anybody who wants to judge the frame before trusting a rate, the scanner's own classification rules are on the methodology page, and what this crawler sends is documented at the bot page. Readers who want the vendor-side version of the question rather than the census should start at how to get cited in Google AI Overviews, and those auditing their markup more broadly will find the structured data reference closer to what they need than this post is.

Measured on 7 October 2026

  • What 1,083 home pages delivered to one client
  • Snippet directives in metas and X-Robots-Tag headers
  • The exact string each of the 132 sent
  • Platform markers on the 129 re-fetched pages

Not measured

  • Whether Google honoured any directive
  • Whether any page reached an AI Mode answer
  • Pages other than the home page
  • Search Console account level opt outs
What this run measured against what it did not, stated so the figures are not read as more than they are.

Written by

Lantad

Published .

Ask how to control what Google's generative surfaces do with a page and the answer almost always comes back as a robots.txt token. Google-Extended is the one people name. It is the wrong answer for AI visibility in Google's own search products, and Google says so on two separate documentation pages that have both been updated this year. The control over AI Mode and AI Overviews is not a crawler token at all. It is a snippet directive, the same unglamorous field that has governed the grey text under a blue link since long before any of this.

Common questions

How do I stop my content appearing in Google AI Mode?

Google's robots meta tag documentation, Last updated 2026-03-24 UTC, states that the nosnippet rule will also prevent the content from being used as a direct input for AI Overviews and AI Mode, and that max-snippet:0 is equivalent to nosnippet. Google also began testing a Search Console setting in June 2026 that excludes a property from AI Overviews and AI Mode, which is an account setting rather than anything on your site. Lantad found 1 of 1,083 readable home pages sending nosnippet and none sending max-snippet:0 on 7 October 2026.

Does blocking Google-Extended remove me from AI Mode?

Google's documentation does not say that it does. The crawler documentation, Last updated 2026-07-14 UTC, scopes Google-Extended to whether crawled content may be used for training Gemini models and for grounding in Gemini Apps and Vertex AI, and states that Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. AI Mode and AI Overviews are Google Search surfaces and that entry names neither of them.

What does max-snippet:-1 actually do?

Nothing restrictive. Google's documentation lists -1 as a special value meaning Google will choose the snippet length that it believes is most effective to help users discover your content and direct users to your site, which is also what happens when no rule is set. Lantad measured 129 of 1,083 readable home pages sending it on 7 October 2026, and found 122 of its 136 occurrences carried the same five-token directive set in a different order, with 89 of the 129 pages carrying a WordPress SEO plugin marker.

Is a snippet directive read if robots.txt blocks the crawler?

No. Google's robots meta tag documentation states that these settings can be read and followed only if crawlers are allowed to access the pages that include these settings, because a meta tag and an X-Robots-Tag header both arrive inside a response the crawler has to be permitted to request. In this run 217 of 1,419 hostnames answered HTTP 403 and 57 answered HTTP 503 to an identified crawler, so whatever directives those 274 pages carry were unreadable from outside.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.