# The AI Overviews opt out does not live on your site

> Google began testing a Search Console setting on 3 June 2026 that excludes a site from AI Overviews and AI Mode. It is an account setting rather than a file, so no external scan of your site can tell you whether it is set.

- Canonical page: https://lantad.co/blog/ai-overviews-opt-out-does-not-live-on-your-site
- This file: https://lantad.co/blog/ai-overviews-opt-out-does-not-live-on-your-site.md
- Last substantive update: 2026-07-28

## Key facts

- **Published:** 2026-07-28
- **Category:** Guides
- **Author:** Lantad
- **Length:** 4951 words
- **Takeaway 1:** Google announced on 3 June 2026 a Search Console control, found under Settings and then Search generative AI, that excludes a property's links and content from AI Overviews, AI Mode and generative AI features in Discover.
- **Takeaway 2:** Google's announcement of 3 June 2026 states that sites which opt out will not receive traffic or impressions from its generative AI features, and that the control will not be used as a ranking signal for search results outside those features.
- **Takeaway 3:** The Search generative AI control is a per property setting inside an authenticated Search Console account, so unlike robots.txt, a meta robots tag or an HTTP header it leaves nothing on the site for an external scanner to fetch.
- **Takeaway 4:** Google's crawler documentation, carrying a last updated date of 14 July 2026, states that Google-Extended does not impact a site's inclusion in Google Search, so disallowing that token has never removed a page from AI Overviews.
- **Takeaway 5:** Lantad has not measured adoption of this control and reports its behaviour from Google's published documentation and the CMA's published conduct requirement as read on 28 July 2026, not from any scan.

## Summary

Until this summer, every control over what an AI system could do with your pages was a file on your server. A robots.txt rule, a meta robots tag, an HTTP response header: different mechanisms with one property in common, which is that anybody could fetch them and read what they said. That property is the reason [AI visibility](https://lantad.co/glossary/ai-visibility) is measurable from outside at all. A scanner needs no account and no permission to tell you what your own site is publishing about itself.

The AI Overviews opt out breaks that pattern. Google began testing a control on 3 June 2026 that removes a site from AI Overviews and AI Mode, and the control does not live on your site. It lives in Search Console, behind a login, as a setting on a verified property. Nothing about the page changes, no header appears, robots.txt is untouched, and a crawler that fetched every URL you own would find no trace of it. This post reports what Google and the UK regulator have published about that control, and then covers the part that matters for anyone measuring what [an AI crawler](https://lantad.co/glossary/ai-crawler) can do with a site: a control that leaves no observable trace is a control no external tool can verify, including this one.

## What Google shipped, and the regulator that required it

The sequence starts with a regulator rather than a product decision. On 3 June 2026 the Competition and Markets Authority published [a press release on securing a fairer deal for publishers](https://www.gov.uk/government/news/cma-secures-fairer-deal-for-publishers-and-improves-google-search-services-in-uk) whose claim about AI features is unusually direct for a document of that kind: in a world first, publishers will now have effective tools to prevent their content being used to power AI features in search, such as AI Overviews. The same release states that Google will now also have to allow publishers to opt out of allowing their content to be used for the fine-tuning of AI models, and that Google is required to make sure publisher content is properly attributed, using clear links, in AI generated search results.

The instrument is a conduct requirement imposed under the UK digital markets competition regime, following the designation of Google as having strategic market status in general search. [The published requirement](https://www.gov.uk/find-digital-markets-measures/google-search-publisher-conduct-requirement) sets out five obligations in flat language. Google must provide publishers with effective controls over the use of their search content in generative AI. It must publish clear, comprehensible and user-friendly information explaining how publishers' search content is used by Google in its generative AI. It must provide publishers with clear and detailed metrics on user engagement with their search content in search generative AI features. It must take reasonable steps to ensure that search content is attributed clearly and accurately in general search, and that end users have a clear means to access that search content. And it must publish clear, comprehensible and user-friendly information explaining its approach to attribution.

On timing the press release is specific. Google will have nine months to implement all changes, but the CMA expects important parts of the controls to become available to publishers well before that deadline. Nine months from 3 June 2026 runs into early March 2027, so full compliance is not due yet and what is visible today is the early part the regulator asked to see first.

Google published on the same day. [Its announcement for website owners](https://blog.google/products-and-platforms/products/search/new-controls-website-owners/) describes beginning to test a new control that lets website owners manage how their links and content appear in generative AI Search features, naming AI Overviews, AI Mode and AI Overviews in Discover. Alongside it, Search Central announced [generative AI performance reports in Search Console](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports), which report impressions, which pages appeared in generative AI features, and the countries they appeared in. Those two shipments map onto the first and third obligations in the list above. The control and the metrics arrived together because they were asked for together.

It is worth being careful about what this does and does not settle. A conduct requirement is a rule about one company's behaviour in one jurisdiction. It is not a specification, it creates no mechanism any crawler reads, and it binds no other vendor. Anyone reading it as the arrival of a general opt out for AI search is reading a great deal into it, in the same way that the [IETF's preference vocabulary](https://lantad.co/blog/the-ai-preferences-standard-you-cannot-deploy-yet) is routinely reported as a deployable standard while the half of it that says where to write a preference down sits expired. What this does settle is narrower and still substantial: for Google's generative surfaces, and initially for a subset of owners in one country, there is now a switch where previously there was only a tradeoff.

## Why no external scan can tell you whether this control is set

Here is the mechanical difference, and it is the whole reason this post exists.

A robots.txt rule is a line in a file at a known path. A noindex or nosnippet instruction is a tag in a document or a field on a response. An llms.txt file, whatever you make of the evidence for it, is a file. Every one of those can be fetched by anyone, from anywhere, with no account, which means an external tool can report what your site says and, more usefully, whether the response that came back matches it. That comparison is [what a scan actually measures](https://lantad.co/methodology): not what you intended, but what a request received.

Google's Search generative AI setting has none of those properties. Google's Search Console help documentation, which is not on a host this site links to and sits at support.google.com/webmasters/answer/16908024, describes a setting found under Settings and then Search generative AI, offering a property three choices: include my site's links and content in Search generative AI features, which is the default, exclude my site's links and content from Search generative AI features, or inherit control from parent. It is set per property, child properties inherit from a parent unless overridden, and the documentation says a change generally takes a few days to take effect.

Read that description again with a scanner's eyes. There is no path to request. There is no header to parse. There is no difference in the bytes your server returns to a crawler between a property that has excluded itself and one that has not. The setting is a record in Google's systems attached to a verified Search Console property, and the only people who can read it are the people who can log in to that property.

The consequence deserves stating without hedging, because it is inconvenient for a product that sells measurement. No external scan can tell you whether a site has opted out of AI Overviews. Lantad cannot. Neither can any other crawler based tool, now or later, because the information is not in the transport. Any dashboard claiming to report a site's AI Overviews eligibility is inferring it from something else, most likely from whether the site currently appears in AI results, which is a different question with many other possible causes.

This is the same discipline as [refusing to print a grade for a page that could not be measured](https://lantad.co/blog/why-we-withhold-a-grade). A tool that reported this setting would be guessing, and a guess presented as a reading is worse than an admitted gap, because the reader cannot tell which of the two they are holding. The [research position published on this site](https://lantad.co/research) is that a figure ships when there is enough measured evidence behind it and not before, and the same logic covers a signal that cannot be observed at all.

There is a second order effect that matters more for diagnosis than the first. Absence from AI Overviews now has one additional possible cause, and it is a cause with no external evidence whatsoever. A page might be missing because a crawler could not read it, which [a rendering check](https://lantad.co/tools/what-gptbot-sees) will show you directly. It might be missing because a layer above your origin refused the request, which is [the case the two layers post sets out](https://lantad.co/blog/two-layers-decide-if-ai-can-read-your-site) using Cloudflare's own announcement. Or it might be missing because somebody with Search Console access changed a setting and nobody wrote it down. The first two are measurable from outside and the third is not, which makes the third the cheapest to rule out and the one to check first, before anybody spends a week on the other two.

## Four Google controls, and none of them substitutes for another

A site owner who wants to change how Google's AI answers use their pages now has four levers, and the expensive mistake is assuming any one of them covers another.

The first is robots.txt addressed to Googlebot. Google's [AI features documentation](https://developers.google.com/search/docs/appearance/ai-features) is explicit that AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot are the control for managing how sites are crawled for Search. Blocking Googlebot removes you from Search, and therefore from the generative features built on top of Search. It works, in the sense that a demolished building no longer has a leaking roof.

The second is the snippet family. The same page names nosnippet, data-nosnippet, max-snippet and noindex as the way to limit the information shown from your pages in Search. These do reach AI features, and the next section is about what they cost.

The third is Google-Extended, and it is the one most often misread. [Google's crawler documentation](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers), carrying a last updated date of 14 July 2026, describes it as a standalone product token publishers use to manage whether content Google crawls may be used for training future generations of Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. The same page states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal in Google Search. AI Overviews and AI Mode are Search. It follows directly that a robots.txt file disallowing Google-Extended has never removed a single page from an AI Overview, and was never documented to.

That is a specific, checkable claim, and it contradicts a lot of published advice. It is also not the only trap in that token. Google-Extended publishes no request user agent of its own, so it can never appear in an access log on any site, which is [set out in full in the post on tokens that never appear in your logs](https://lantad.co/blog/the-crawler-tokens-that-never-appear-in-your-logs). A control that governs a narrower surface than people believe, and that leaves no trace when it is working, is an unfortunate combination of properties.

The fourth is the new Search Console setting, covering AI Overviews, AI Mode and generative AI features in Discover, and nothing beyond them. Per Google's help documentation it does not stop content being used for AI model training, which remains Google-Extended's job, and it does not affect Merchant Center or Google Ads.

Laid side by side the division is simple enough to hold in your head. Crawling for Search is robots.txt and Googlebot. What may be displayed from a page in Search, including inside AI answers, is the snippet family. Gemini training and grounding outside Search is Google-Extended. Presence in AI Overviews, AI Mode and Discover's generative features is the Search Console setting. Four levers, four surfaces, no useful overlap, and only three of the four leave anything on your site to look at.

Anyone working seriously on [generative engine optimisation](https://lantad.co/glossary/geo) should be able to name which lever governs which surface without checking, because the failure is silent in both directions. Pull the wrong one and you remain in the feature you meant to leave, with nothing to tell you. Pull the right one and you leave a surface you may have wanted to be in, which is a genuine commercial decision rather than a technicality: the [notes on getting cited in AI Overviews](https://lantad.co/how-to-get-cited/google-ai-overviews) exist because appearing there is worth something to plenty of sites.

## The old lever cost you your search snippet, and the new one does not

To see why a setting in a dashboard is worth anything at all, it helps to look at what the alternative asked you to give up.

Before June 2026 the documented way to keep your text out of an AI Overview was the snippet family, and the snippet family does not distinguish between surfaces. nosnippet does not mean no AI snippet. It means no snippet: the same directive that keeps your sentences out of a generated answer keeps them out of the ordinary result underneath it. max-snippet is a length limit applied wherever Google shows text from your page. data-nosnippet is finer, marking a region of a document rather than the whole of it, but it is still expressed in the same currency. You bought your way out of AI answers by paying with your presence in classic results, and for most publishers that price was higher than the problem.

That is what makes the new control interesting, and it is the part worth reading carefully in [Google's announcement](https://blog.google/products-and-platforms/products/search/new-controls-website-owners/). Two statements sit next to each other there. Sites that opt out will not receive traffic or impressions from its generative AI features, which is the cost, stated plainly and without softening. And the control will not be used as a ranking signal for search results outside of these generative AI Search features, which is the thing that was not previously on offer. Google's help documentation goes further in the same direction, saying the exclusion does not affect regular Search rankings or indexing and does not affect Merchant Center or Google Ads.

Read together, those describe a decoupling. The question moved from "how much of Search am I willing to lose to stay out of AI answers" to "do I want to be in AI answers, yes or no". Those are different questions and only the second one is answerable on its merits.

Whether the answer should be no is a separate argument and not one this post can settle for anybody. It turns on numbers each publisher has and this site does not: what proportion of your traffic arrives through generative surfaces, what a citation without a click is worth to you, and whether your content is the sort a model can summarise away or the sort people come to the source for. What can be said from the documentation is that the decision is now available at a price that is described, which was not true two months ago, and that the [AI features documentation](https://developers.google.com/search/docs/appearance/ai-features) still described only the snippet route in the control section that was live when this post was written, carrying a last updated date of 10 December 2025. Documentation lags shipping; if you go looking for the new control on that page today you may not find it, and that is not evidence it does not exist.

One caution about the opposite direction, since most readers of this site want more visibility rather than less. Nothing in this control helps you appear in AI Overviews. It is an exclusion switch with a default of inclusion, so leaving it alone is already the maximally visible position, and there is no setting that improves your odds. The things that improve your odds are the same unglamorous ones as before: pages a crawler can actually read, which is what [prose parity](https://lantad.co/glossary/prose-parity) measures, and enough structure for a machine to work out what the page is about, which is where [structured data](https://lantad.co/glossary/structured-data) earns its place. A switch that governs whether you are eligible is not a switch that makes you worth citing, and it would be a poor trade to spend a month on the first while ignoring the second. For the difference between ranking and being quoted, [answer engine optimisation](https://lantad.co/glossary/aeo) is the term that names it.

## Availability depends on your jurisdiction, not on your site

There is one more property of this control that has no precedent in the file based world, and it is easy to miss because it is stated in a single clause.

Google's announcement describes beginning to roll these features out to a subset of website owners in the UK, allowing for thorough testing before rolling them out to website owners globally. So whether the control exists for you is currently determined by where you are and whether you are in a test group, and not by anything about your pages. Two identical sites, same stack, same content, same robots.txt, can have different options available in their Search Console accounts. That has never been true of robots.txt, which every crawler on the open web reads the same way from every country, under [the specification published as RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html).

This is worth naming clearly, because a jurisdictional control is a different kind of object from a technical one, and reasoning about it with technical instincts produces errors. A specification either exists or it does not, and if it exists you can implement it. A regulated control exists for the people a regulator has jurisdiction over, arrives on the schedule the regulated company chooses within the deadline it was given, and may or may not generalise. The CMA's press release said Google has nine months to implement all changes, with important parts expected well before that. Nothing in it obliges Google to give the same control to a site owner in Ohio or Osaka, and Google's own wording, that global rollout follows testing, is a statement of intent rather than a commitment to a date.

The practical consequence for a reader outside the UK is that this control is a thing to know about rather than a thing to use, for now. The practical consequence for everyone is subtler. AI visibility has acquired a second axis: not only what your site does, which is what tooling can measure, but which controls your account is offered, which it cannot. Adding that axis to a mental model that only had the first one is the point of this post.

The companion shipment is easier to act on wherever you are, because reporting was the other half of what the regulator asked for. [The generative AI performance reports](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports) give impressions in generative AI features, the pages that appeared, and the countries they appeared in. That is first party data about a surface that was previously almost opaque, and it comes from Google rather than from an inference. It is worth saying plainly that this is better evidence about your presence in AI Overviews than any third party tool can produce, this one included, for the same reason the setting is invisible: the data lives in Google's systems and is released to the verified owner.

That asymmetry is not unique to Google, and it is the general shape of measuring anything in this market from outside. A per crawler figure counts requests that claimed to be a given crawler, which is why [a user agent is a claim rather than an identity](https://lantad.co/blog/a-user-agent-is-a-claim-not-an-identity) and why verification costs a reverse DNS or published IP range check. Platform behaviour differs enough that the guidance splits by platform, which is why the notes on being cited by [ChatGPT](https://lantad.co/how-to-get-cited/chatgpt) and by [Perplexity](https://lantad.co/how-to-get-cited/perplexity) are separate pages rather than one. External measurement is genuinely useful and genuinely bounded, and the boundary moved this summer.

## What to check on your own site this week

None of the above changes what a page has to do to be worth citing. It changes the order in which you should rule things out, and it adds one check that no tool can do for you.

Start with the one that is now first, because it is free and it is invisible to everyone else. Open Search Console, go to Settings, and look for Search generative AI. If the section is not there, your property is not in the rollout and there is nothing to do. If it is there, read which option is selected, and read it on every property that matters, including child properties, since those inherit from a parent unless somebody has overridden them. The reason to look even when you are confident is that this is the only AI visibility setting on your estate that nobody can audit from outside, which means the usual safety net of somebody eventually noticing does not exist.

Then check the things that are observable, in the order that a request meets them. First, whether your robots.txt admits the crawlers you intend it to, per token rather than as one verdict, since the common failure is a file that admits one fleet and refuses another with nobody intending it: [the AI crawler check](https://lantad.co/tools/ai-crawlers) tests each token separately, and [the robots.txt tester](https://lantad.co/tools/robots-txt-tester) answers the same question for a specific path and agent. While you are in that file, confirm what you believe about Google-Extended matches [what Google documents it to govern](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers), because a rule written on the belief that it governs AI Overviews is a rule doing something other than what its author intended.

Second, whether the response a crawler receives is the same one a browser receives. This is where most sites actually lose, and it has nothing to do with permission: a page that renders its content client side can return an empty shell to a fetcher that does not execute JavaScript, and [what GPTBot sees](https://lantad.co/tools/what-gptbot-sees) shows the difference on a URL you choose. If that check comes back thin, the fix is stack specific rather than general, which is why the guidance splits into per stack pages for [Next.js](https://lantad.co/fix/nextjs), [React](https://lantad.co/fix/react) and [Shopify](https://lantad.co/fix/shopify) rather than one page of advice that fits nobody.

Third, whether anything above your origin is answering on its behalf. A CDN, a firewall or a bot management rule can refuse a crawler that your robots.txt welcomes, and the file will never mention it because the file is not the layer that answers. That is the whole argument of [the two layers post](https://lantad.co/blog/two-layers-decide-if-ai-can-read-your-site), and it is unchanged by anything Google shipped in June.

Fourth, and only once the first three are clean, the content questions. Whether a machine can tell what the page is about, which is what [entity confidence](https://lantad.co/glossary/entity-confidence) is trying to capture and which [the five signals post](https://lantad.co/blog/five-signals-that-tell-ai-who-you-are) breaks down signal by signal. Whether the claims on the page are the sort a model can attribute to you. And whether any of the files you have been told to add are doing anything, a question worth asking sceptically given [the published evidence on llms.txt](https://lantad.co/blog/what-the-evidence-says-about-llms-txt), where Ahrefs measured 137,210 domains in May 2026 and found 97 percent of valid files were never fetched at all.

The order matters more than the list. A setting nobody can see beats a rendering bug for diagnostic priority precisely because it is cheap to check and impossible to detect any other way, and rendering beats content because content nobody can read scores nothing whatever it says. For what this scanner sends and how to refuse it, [LantadBot's conduct policy](https://lantad.co/bot) documents the token, on the principle that a tool asking you to audit crawler behaviour should be auditable itself. And if any of the above turns out to need re-measuring after a change, [how a page is scored](https://lantad.co/methodology) sets out what is weighted and what is only reported.

## Questions and answers

**Does blocking Google-Extended remove my site from AI Overviews?**

No. Google's crawler documentation, carrying a last updated date of 14 July 2026, states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal in Google Search. AI Overviews and AI Mode are Search features, so a robots.txt rule disallowing Google-Extended does not remove pages from them. What that token governs is whether crawled content may be used for training future generations of Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI.

**Can a scanner tell me whether a site has opted out of AI Overviews?**

No, and neither can Lantad. Google's control is a setting on a verified Search Console property, found under Settings and then Search generative AI. It changes nothing about the bytes a server returns, adds no header, and touches no file, so there is no request an external tool can make that would reveal it. Only somebody with access to that Search Console property can read it. Any tool reporting AI Overviews eligibility is inferring it from something else, most often from current presence in AI results.

**What does opting out of Search generative AI actually cost?**

Google's announcement of 3 June 2026 states that sites which opt out will not receive traffic or impressions from its generative AI features, and that the control will not be used as a ranking signal for search results outside of those features. Google's help documentation adds that the exclusion does not affect regular Search rankings or indexing, Merchant Center or Google Ads, and does not stop content being used for AI model training, which is governed separately by Google-Extended.

**Is this control available to everyone?**

Not yet. Google's announcement of 3 June 2026 describes beginning to roll the features out to a subset of website owners in the UK before rolling them out to website owners globally, and no date has been published for the global rollout. The control follows a conduct requirement the CMA imposed on 3 June 2026, under which Google has nine months to implement all changes, with the CMA expecting important parts to become available well before that deadline.

---

Lantad measures whether AI crawlers can actually read a page: it fetches as a non-rendering
crawler, renders as a browser, and reports the gap. Free scan, one URL, no signup.

Method and weights: https://lantad.co/methodology | All pages as markdown: https://lantad.co/md | Crawler policy: https://lantad.co/bot
