Blog / AI Overviews retrieved less from sites blocking Google-Extended

AI Overviews retrieved less from sites blocking Google-Extended

A SIGIR 2026 study of 11,500 queries collected on 7 and 8 December 2025 found Google AI Overviews significantly less likely than Google Search to retrieve content from sites that block Google-Extended, the token Google documents as not affecting Search. Gemini never cited the study's 21 blocking publishers at all.

In short

  • How Generative AI Disrupts Search, posted to arXiv as 2604.27790 on 30 April 2026 and accepted to SIGIR 2026, found Google AI Overviews and Gemini 2.5 Flash both significantly less likely than Google Search to retrieve content from sites that block Google-Extended in robots.txt, despite AI Overviews having access to that content through Googlebot.
  • The study identified 21 popular publishers, including NYTimes, BBC, Reuters and Nature, that Google Search and AI Overviews each retrieved for at least 20 unique queries and Gemini never cited once; all 21 block Google-Extended in their robots.txt files.
  • Google's crawler documentation, last updated 14 July 2026, states Google-Extended controls Gemini training and grounding and does not impact a site's inclusion in Google Search, and it does not mention AI Overviews.
  • Across 11,500 queries collected on 7 and 8 December 2025, the average Jaccard overlap between the sources AI Overviews cite and the Google Search results below them was 0.18, and AI Overviews appeared on 51.5 percent of the 5,000 real user queries in the benchmark.
  • Lantad did not run this study and measured none of these figures, which are reported from the paper as read on 6 August 2026; Google-Extended is one of the 15 crawler tokens in Lantad's registry, which is a setting, not a finding.

Google-Extended is the opt out that is supposed to cost nothing inside Search. Google's crawler documentation describes a token that controls whether content trains future Gemini models and grounds Gemini Apps and Vertex AI, and states that it does not impact a site's inclusion in Google Search and is not a ranking signal. A publisher reading that page can decline to feed Gemini while keeping everything Search sends, and much of the news industry has taken the trade on those terms.

A study accepted to SIGIR 2026 measured the terms. Across 11,500 queries put to Google Search, AI Overviews and Gemini 2.5 Flash on 7 and 8 December 2025, sites blocking Google-Extended were significantly less likely to be retrieved by AI Overviews than by Google Search for the same queries, despite AI Overviews technically having access to their content. And the 21 popular publishers that Gemini never cited at all turned out, on inspection, to share one property: every one of them blocks Google-Extended. Lantad measured none of what follows. This is a reading of one paper, opened on 6 August 2026, about a robots.txt decision most sites made without expecting any consequence inside Search.

  • Googlebot Access unchanged Google-Extended does not gate crawling. Googlebot fetches and indexes blocking sites exactly as before.
  • Google Search Comparison baseline The traditional results page is the engine the study's regression measures the other two against.
  • AI Overviews Retrieved less Significantly less likely than Search to retrieve content from blocking sites, despite technically having access to it.
  • Gemini 2.5 Flash 0 of 21 publishers Never cited any of the 21 blocking publishers across the benchmark. The grounding opt out behaved as documented.
How each Google surface treated sites blocking Google-Extended, from arXiv 2604.27790, data collected 7 and 8 December 2025. Reported from the paper as read on 6 August 2026, not measured by Lantad.

What the study measured, and when

The paper is How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews, posted to arXiv as 2604.27790 on 30 April 2026 and accepted to SIGIR 2026, written by six authors at the New Jersey Institute of Technology, Nanyang Technological University and Indiana University Bloomington. Its object of study is the same query asked three ways: the traditional Google Search results page, the AI Overview generated above it, and Gemini 2.5 Flash answering as a standalone assistant. For each of 11,500 benchmark queries the authors recorded which sources each of the three systems retrieved, then compared the lists.

The benchmark is nine query sets rather than one. The largest is ORCAS, 5,000 real user queries drawn from a public search log, and the paper treats it as the reference point for how the feature behaves on queries people actually type. Around it sit three Amazon retail sets totalling 1,500 shopping queries, 1,000 debate style questions, 1,000 ELI5 explanation questions, 1,000 localized queries and 2,000 built on Natural Questions in question and keyword form; three of the nine sets are synthetic variants the authors generated with Gemini 2.5 Flash, which the paper discloses. The collection is a snapshot: all Search, AI Overview and Gemini responses to benchmark queries were collected on 7 and 8 December 2025. That date belongs next to every figure in this post. These systems change monthly, and a December 2025 measurement describes December 2025.

The first headline number is how often the generative layer appears at all. AI Overviews were generated on 65.6 percent of all benchmark queries, and on 51.5 percent of the ORCAS real user set, and the authors write that they consider the ORCAS rate the most representative of how often a real user query results in a generated AI Overview. Half of real queries in this sample opened on a generated answer, and everything else in the paper is about what that answer chooses to stand on.

Two of the study's boundaries matter for everything below. The retrieval comparisons are relative: the analysis dataset contains only sources that at least one of the three engines retrieved, so every finding reads as one engine against another rather than as an absolute rate. And the collection window is two days rather than a monitoring period, so the paper establishes what the systems did that week, not what they do this one.

Query setQueries
ORCAS, real user queries from a public search log5,000
Amazon retail, three sets including synthetic variants1,500
Debate questions1,000
ELI5 explanation questions1,000
Localized queries1,000
Natural Questions, question form1,000
Natural Questions, keyword form1,000
Composition of the 11,500 query benchmark in arXiv 2604.27790, counts as published in the paper's description of its query sets. Reported from the paper, not measured by Lantad.

What Google says Google-Extended controls

Google-Extended is not a crawler. Google's crawler documentation, carrying Last updated 2026-07-14 UTC and read on 6 August 2026, lists it as a standalone product token that manages whether content is used to train future generations of Gemini models and for grounding in Gemini Apps and Vertex AI, and states directly that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal there. There is no Google-Extended user agent. The fetching is done by Googlebot, and a disallow under the token changes what Google may do with a page rather than whether the page is fetched, a mechanism set out here before in the crawler tokens that never appear in your logs, along with why a log search for the token returns zero on every site in the world.

The documentation for the answer side says the control lives elsewhere. Google's AI features page tells site owners that robots.txt directives for Googlebot are the access control for AI features in Search, points to the preview controls, nosnippet, data-nosnippet, max-snippet and noindex, for limiting what appears in them, and mentions Google-Extended only as the way to limit AI training and grounding in some of Google's other systems. Read together, the two pages describe a clean split: Googlebot and the preview controls govern Search and its AI features, Google-Extended governs Gemini, and the two do not touch. A third control exists on neither page, the Search Console setting that excludes a property from AI Overviews and AI Mode, which is an account setting rather than a file and was covered in the AI Overviews opt out does not live on your site.

Mechanically the token behaves like any other group name. Under RFC 9309 a crawler selects the group matching its product token, so a group headed User-agent: Google-Extended binds only a client answering to that name and the wildcard group continues to decide everyone else. Google-Extended is one of the 15 tokens in Lantad's crawler registry, and that is a setting rather than a finding: someone chose to track it because publishers write rules for it at scale.

What the paper adds is a measurement of the clean split. On the Gemini side, the split holds perfectly. On the AI Overviews side, the correlation in the next two sections is not supposed to exist.

What the documentation says

  • Controls Gemini training and grounding
  • Does not impact inclusion in Google Search
  • Not used as a ranking signal
  • AI features access is governed by Googlebot
  • Crawler page last updated 14 July 2026

What the study measured

  • AI Overviews significantly less likely than Search to retrieve blockers
  • Gemini cited 0 of the 21 blocking publishers
  • Gap present despite Googlebot access
  • Data collected 7 and 8 December 2025
The documented behaviour of Google-Extended against what arXiv 2604.27790 measured in December 2025. The left column is Google's documentation as read on 6 August 2026; the right is the paper's findings, not a Lantad measurement.

The 21 publishers Gemini never cited

The paper's Table 3 was found backwards, and the method is worth restating because it is what makes the list persuasive. The authors did not start from robots.txt files. They started from retrieval counts, and noticed a set of popular publishers that Google Search and AI Overviews each retrieved for at least 20 unique queries and that Gemini cited zero times across the whole benchmark. Twenty one publishers met that description. Upon further investigation, the authors write, all of these websites block the Google-Extended bot in their robots.txt files.

The names are not marginal sites. Ranked by Tranco popularity, the list runs NYTimes, ESPN, Genius, CNN, Business Insider, National Geographic, BBC, CNBC, The Conversation, ScienceDirect, NPR, U.S. News & World Report, Reuters, WIRED, Scientific American, Wiley, USA Today, Consumer Reports, Nature, NBC News and STAT. That composition is what the population statistics predict rather than a coincidence the paper manufactured: the robots.txt study covered here in the sites that block AI crawlers are the ones with editors put Google-Extended disallows at 43.51 percent of reputable news sites in its May 2025 reading, so a list of high credibility publishers blocking the token is exactly the expected shape.

For Gemini, this is the opt out working precisely as documented. Grounding declines content whose owner reserved it, and in this sample it declines it completely: not reduced citation but zero citation, across the benchmark, for all 21 publishers. A publisher who wrote that disallow line to keep Gemini from leaning on its journalism got, on this evidence, exactly what the line promised.

The same table is what makes the AI Overviews finding a reduction rather than an absence. The 21 publishers qualified for the list because AI Overviews did retrieve them, for at least 20 unique queries each. Blocking Google-Extended does not remove a site from AI Overviews, and nothing in the paper says otherwise. What the full text reports is quieter and stranger: for the same queries, the AI Overview reaches for those publishers significantly less often than the results page directly beneath it does.

PublisherGemini citations in the benchmark
NYTimes0
ESPN0
Genius0
CNN0
Business Insider0
National Geographic0
BBC0
CNBC0
The Conversation0
ScienceDirect0
NPR0
U.S. News & World Report0
Reuters0
WIRED0
Scientific American0
Wiley0
USA Today0
Consumer Reports0
Nature0
NBC News0
STAT0
The 21 publishers in Table 3 of arXiv 2604.27790, ranked by Tranco popularity: each was retrieved by Google Search and by AI Overviews for at least 20 unique queries in the benchmark, and never cited by Gemini 2.5 Flash. All 21 block Google-Extended in robots.txt. Reported from the paper, not measured by Lantad.

Retrieved less despite access: what the regression can and cannot say

The blocking result comes from the paper's source level regression, and its construction decides what the result can mean. The authors build a dataset with every unique pair of source and query, and mark whether each of the three engines cited that source for that query. Because a source only enters the dataset if at least one engine retrieved it, the authors state the boundary themselves: the results can only be interpreted in comparison to the other search engines. This is a relative measurement by design. It cannot say AI Overviews cite blockers rarely. It says AI Overviews cite blockers less than Google Search does, on the same queries, at the same moment.

On that dataset they fit a linear probability model with query fixed effects and clustered standard errors, comparing each generative engine against traditional search, with site characteristics as regressors: the site's Tranco popularity rank in bins, its domain category, whether it is an .edu, and whether it blocks Google-Extended. The blocking coefficient is negative and statistically significant for both generative systems, and the paper's summary carries the part that matters: both Gemini and AI Overviews are significantly less likely to retrieve content from websites blocking Google-Extended, despite AI Overviews technically having access to this content.

Despite having access is the phrase to sit with. An AI crawler that is refused at the network or in the file never sees the page, and reduced citation follows trivially; that is not this. Googlebot fetches these publishers, indexes them and ranks them, and the blue links on the same results page retrieve them at the full rate. Whatever produces the gap operates after access, inside retrieval, where no publisher can observe it. As covered in AI Overviews and the one crawler our scanner does not model, the AI Overview pipeline begins with Googlebot, and Googlebot's treatment of these sites is unchanged.

What the regression does not establish is mechanism, and the paper does not claim one. Popularity and category controls rule out the crudest confounders, but the 21 blockers are also, almost uniformly, publishers with licensing positions, paywalls and legal teams, and observational data cannot separate the token from everything else those publishers do. Whether Google consults Google-Extended when assembling an AI Overview, whether AI Overview grounding shares retrieval infrastructure with Gemini, or whether something else entirely produces the gap is not in the paper. The correlation is. On December 2025 data, it is the only public measurement of this question at all.

The path from a Google-Extended disallow to each of the three surfaces, as measured in arXiv 2604.27790. A restatement of the paper's finding, not a mechanism: the paper does not identify why AI Overview retrieval falls.

Three engines, three different source lists

The blocking finding sits inside a larger measurement of divergence, and the divergence recalibrates what any single citation check can tell you. For each query the authors computed the overlap between the sources each pair of engines retrieved, as Jaccard similarity, the size of the intersection over the size of the union. The averages are low everywhere: 0.18 between AI Overviews and the results page they sit above, 0.16 between the results page and Gemini, and 0.11 between AI Overviews and Gemini. The paper puts the first number in words: on average, only 18 percent of the sources returned by either the AI Overview or the traditional results page will be retrieved by both.

Read that against the habit of treating Google rank as a proxy for AI visibility. The two lists share less than a fifth of their entries, on the same query, from the same company, on the same page of results. It is the same direction of finding as the citation study covered in nearly a third of AI Overview citations are not on page one, measured a different way, and it extends to the assistant: Gemini's sources overlap the results page slightly less than the AI Overview's do, and a tool that measures Gemini is measuring a third surface again, which is the boundary drawn in a Gemini citation tool is not an AI Overview tool.

The paper also measured how stable each surface is, and the generative one moves more. Repeating the same query from the same device and location, AI Overviews scored a rank biased overlap of 0.69 against 0.86 for the results page; from a different device the pair is 0.53 against 0.79. Cosmetic edits to a query, changes that leave the meaning intact, drop AI Overview overlap to 0.49 while the results page holds 0.74 and Gemini 0.5. For anyone measuring AI visibility, that instability is the sampling problem quantified before in how many prompts an AI visibility measurement needs: a single query run once tells you what one roll of the dice did, and a December 2025 AI Overview rolled more dice than the results page under it.

One more compositional finding belongs in the record. The paper reports that AI Overviews favor Google websites, the google.com and youtube.com domains, relative to traditional search, and that Gemini also favors YouTube with a smaller absolute difference. Retrieval that leans toward the answer engine's own properties compresses the space left for everyone else before any per site factor applies.

  • AI Overviews and Google Search 0.18 Only 18 percent of combined sources appear in both
  • Google Search and Gemini 0.16 The assistant overlaps the results page even less
  • AI Overviews and Gemini 0.11 The two generative surfaces agree least of all
Average Jaccard similarity of retrieved sources between each pair of engines across the benchmark, from Table 2 of arXiv 2604.27790, collected 7 and 8 December 2025. A value of 1 would mean identical source lists. Reported from the paper, not measured by Lantad.

What this changes about the robots.txt decision

Nothing in this study measures your site, and the honest use of a population regression is to go and check the one case you control. The first check is knowing what your file actually says, which is less obvious than it sounds: robots.txt is evaluated per crawler token, group by group, and the outcome for Google-Extended can differ from the outcome for GPTBot inside the same file. Resolving a real file against a specific token is what the robots.txt tester does, and the AI crawler reference enumerates the tokens worth resolving, Google-Extended among them. Whatever this paper's gap turns out to be, a site that does not know whether it blocks the token is not in a position to reason about it.

The second is to hold the two goals apart, because the study is evidence they are coupled in one direction only. If the goal is keeping content out of Gemini training and grounding, the paper says the token delivered, completely, for all 21 publishers in its table. If the goal is being cited in AI Overviews, the paper says the same disallow line correlated with fewer citations than Search rank implies, and Google's documentation says it should not. Those two goals usually belong to different people at a publisher, and one robots.txt line serves the first while possibly taxing the second. A decision that coupled should be made once, deliberately, and re-measured, rather than inherited from a blocklist pasted in during 2024.

The third is to spend effort where the mechanism is known. Whatever explains the blocking gap, the paper's larger finding is that AI Overviews retrieve their own list, weakly overlapping rank and unstable across runs. What a site owner controls is whether the page is readable when retrieval does arrive: whether the text survives a fetch without JavaScript, which What GPTBot sees shows directly, and whether the served page and the rendered page carry the same words, the failure prose parity is named for. Google's own guidance, read here in Google's generative AI guide names five GEO tactics you can ignore, says the rest is ordinary technical health, and this paper gives no reason to disbelieve it.

Where this fits a generative engine optimization programme is as calibration: one December 2025 measurement of how three Google surfaces actually retrieved, and one documented opt out with an undocumented correlation attached. Lantad's scanner reports the file layer and the readability layer for one domain, with its boundaries written in the methodology. An AI visibility score built on those layers is a statement about access and readability, never a promise about retrieval, and this paper is a measured example of why that boundary sits where it does.

  • Resolve robots.txt for Google-Extended by name The token binds its own group. A wildcard disallow does not opt you out of Gemini training, and a Google-Extended group changes nothing for any other crawler.
  • Expect no log evidence either way Google-Extended sends no user agent of its own. Enforcement rides on ordinary Googlebot requests, so an access log can neither confirm nor deny the policy.
  • Name the goal the disallow serves Keeping content out of Gemini grounding worked completely for the study's 21 publishers. If the goal is AI Overview citations, the same line correlated with fewer.
  • Re-measure when either side moves Google's crawler page was last updated 14 July 2026 and the study is a December 2025 snapshot. The documentation and the behaviour can each change without the other.
Four checks that follow from the study for a site owner deciding what to do with Google-Extended. A description of what each check answers, not a measurement of any site.

Related

Common questions

Does blocking Google-Extended remove my site from Google AI Overviews?

No. In the SIGIR 2026 study, the 21 publishers that block Google-Extended were each retrieved by AI Overviews for at least 20 unique queries; that is how they entered the analysis at all. The finding is relative: on data collected 7 and 8 December 2025, AI Overviews were significantly less likely than Google Search to retrieve content from blocking sites on the same queries. The study is correlational, does not identify a mechanism, and Google's documentation states the token does not affect Search.

What does Google-Extended actually control?

Google's crawler documentation, last updated 14 July 2026, describes Google-Extended as a standalone robots.txt product token that manages whether content is used to train future generations of Gemini models and for grounding in Gemini Apps and Vertex AI. It is not a crawler: it sends no requests of its own, and Googlebot does the fetching regardless of the token. The same page states it does not impact a site's inclusion in Google Search and is not used as a ranking signal.

Do AI Overviews cite the same pages that rank in Google Search?

Mostly not. Across the study's 11,500 queries collected on 7 and 8 December 2025, the average Jaccard similarity between the sources an AI Overview cites and the results page beneath it was 0.18, meaning only about 18 percent of the combined sources appear in both lists. Overlap between the two generative surfaces was lower still: 0.11 between AI Overviews and Gemini, against 0.16 between Gemini and the results page.

How often do AI Overviews appear on real queries?

The study measured AI Overviews being generated on 51.5 percent of its 5,000 ORCAS real user queries, the set its authors consider most representative of real usage, and on 65.6 percent of all 11,500 benchmark queries, collected on 7 and 8 December 2025. The rate varies by query category, so shopping, debate and explanation queries do not trigger the feature equally.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.