Blog / Nearly a third of AI Overview citations are not on page one

Nearly a third of AI Overview citations are not on page one

A Washington University in St. Louis team issued 55,393 trending queries over 40 days and found that 29.8 percent of the domains an AI Overview cites appear nowhere on the first page of results for the same query.

In short

  • Researchers at Washington University in St. Louis issued 55,393 trending queries between 13 March and 21 April 2026 and recorded an AI Overview on 7,583 of them, an activation rate of 13.7 percent.
  • In that study, 29.8 percent of the domains cited by an AI Overview appeared nowhere on the first page of results for the same query, and Google's own AI features documentation describes a query fan-out technique that issues multiple related searches to assemble a response.
  • Question-form queries triggered an AI Overview 64.7 percent of the time against 9.5 percent for other phrasings, a 6.8 times difference measured across the same 40 day window.
  • Of 98,020 atomic claims extracted from those AI Overviews and checked against the pages they cite, 11.03 percent were unsupported: 6.98 percent omitted, 2.66 percent contradicted and 1.39 percent ambiguous.
  • Lantad did not run this study and does not measure AI Overviews: every figure here is reported from the paper, arXiv:2605.14021, posted 13 May 2026.

A team at Washington University in St. Louis published a measurement of Google AI Overviews on 13 May 2026, and one number in it should change how a site owner thinks about AI Overview citations. Across 55,393 trending queries issued over 40 days, 29.8 percent of the domains an AI Overview cited did not appear anywhere on the first page of results for the same query. Ranking on page one is not what puts a page into the answer above it, and it is not a prerequisite either.

Lantad did not run this study and cannot run one like it: this scanner does not capture Google results pages and does not model Googlebot, which our own audit of that blind spot sets out line by line against our own source. What follows reports the paper's figures with their sample sizes and dates attached, and then covers the part that belongs to a scanner: which of the things this study describes you can check on your own site, and which of them nobody outside Google can check at all.

  • Also in the top 5 results 25%
  • Also in the top 10 results 41.4%
  • Anywhere on the first page 70.2%
  • Nowhere on the first page 29.8%
Overlap between AI Overview reference domains and the first page results for the same query, averaged across 7,583 overviews. Xu, Iqbal and Montgomery, 13 March to 21 April 2026, arXiv:2605.14021.

What the study measured, and over what sample

The paper is Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact by Haofei Xu, Umar Iqbal and Jacob M. Montgomery of Washington University in St. Louis, posted to arXiv as 2605.14021v1 under cs.CY on 13 May 2026. The collection window ran 40 days, from 13 March to 21 April 2026.

The method decides what the numbers are evidence of, so it is worth stating before any of them. The authors navigated the US localised Google Trends dashboard every 24 hours, exported the trending queries for 19 topical categories, then issued each query to Google in a fresh, stateless browser instance driven by a Puppeteer crawler on AWS Lambda. The crawler waited for the networkidle0 event, meaning no outstanding network requests for 500 milliseconds, before deciding whether an AI Overview was present on the rendered page. For each result it captured the overview, its embedded reference citations, the first page results shown alongside it, and the content and advertising of every cited page.

Two consequences follow, and the paper is direct about both. The sample is trending queries rather than a random sample of everything people search, so it carries whatever the public was interested in during those 40 days: Sports alone accounts for 51.4 percent of the corpus and Entertainment for another 14.9 percent. And the measurement is of a rendered results page as a browser receives it, which is the opposite end of the pipe from what an AI crawler access check looks at. The two answer different questions and neither substitutes for the other.

  • Activation 55,393 trending queries issued over 40 days across 19 Google Trends categories. 7,583 returned an AI Overview.
  • Source selection 61,212 cited reference URLs from 7,479 unique hostnames, compared against the first page results captured on the same page.
  • Claim fidelity 98,020 atomic claims from 7,491 verifiable overviews, each checked against the pages that overview cites.
  • Publisher exposure Advertising markup on every cited page the pipeline could scrape, plus sponsored slots on the results page itself.
The four dimensions the paper measures and the sample behind each. Xu, Iqbal and Montgomery, arXiv:2605.14021, posted 13 May 2026.

How often does an AI Overview actually appear?

Of the 55,393 queries, 7,583 returned an AI Overview. That is 13.7 percent, and it is the first figure to hold onto, because a great deal of writing about generative engine optimisation implies the surface is everywhere. On this sample, in these categories, during this window, it was on roughly one search in seven.

Google says something compatible in its own words. Its AI features documentation, carrying a last updated date of 10 December 2025, states that AI Overviews are only shown when its systems determine that it is additive to classic Search, and as such, often do not trigger. A measured 13.7 percent and a vendor sentence saying the feature often does not fire are the same claim approached from two directions, which is about as much corroboration as this field usually offers.

The aggregate hides the structure that is actually useful. The study classified a query as question form when its leading whole word is one of 15 interrogatives: who, what, where, when, why, how, which, is, are, was, can, do, does, did and has. Question form queries triggered an overview 64.7 percent of the time. Everything else triggered one 9.5 percent of the time. That is a 6.8 times difference within one corpus over one window.

By category the spread is wider still, a factor of 13. Hobbies and Leisure sat at 46.1 percent and Science at 39.9 percent; Beauty and Fashion at 3.5 percent and Climate at 7.4 percent. Politics came in at 7.5 percent and Law and Government at 9.6 percent, which the authors read as consistent with Google exercising caution on sensitive topics, while noting that Health at 26.6 percent shows the caution is not applied evenly across everything sensitive.

Be careful what you carry away from that. The phrasing that predicts an overview is the phrasing of the searcher's query, not the phrasing of your headings. Nothing in the paper says a page written as a question is more likely to be cited. What the split tells you is where the surface exists at all, which is a targeting question rather than an on-page one, and a reason to know which of your terms get asked as questions before spending anything on answer engine optimisation.

  • Hobbies and Leisure 46.1% 1,185 queries
  • Science 39.9% 1,560 queries
  • Health 26.6% 278 queries
  • All 19 categories 13.7% 55,393 queries
  • Law and Government 9.6% 2,924 queries
  • Politics 7.5% 2,166 queries
  • Beauty and Fashion 3.5% 142 queries
AI Overview activation rate by Google Trends category, selected rows with the query count behind each. Xu, Iqbal and Montgomery, 55,393 queries, 13 March to 21 April 2026.

Where AI Overview citations come from, if not page one

Across the 7,583 overviews, Google cited 61,212 reference URLs drawn from 7,479 unique hostnames, a median of 8 references per overview with a range of 1 to 32. Those references are not a re-ranking of the results underneath them. Averaged across the overviews, 25.0 percent of reference domains overlapped with the top five first page results, 41.4 percent with the top ten, and 70.2 percent with the full first page. The remaining 29.8 percent of reference domains appeared nowhere on the first page for that query. At URL level the same gap reads 28.5 percent, 17,451 of the 61,206 references with a parseable host.

The authors' reading is that the AI Overview system draws on a source pool and a prioritisation distinct from Google's own ranking algorithm. Google's documentation describes a mechanism that fits without contradicting them: both AI Overviews and AI Mode may use a query fan-out technique, issuing multiple related searches across subtopics and data sources, which the same page says allows Google to display a wider and more diverse set of helpful links associated with the response than a classic web search shows. A fan-out across subtopics will surface pages that the original query never ranked.

The study also asked whether the two pools differ in quality, scoring domains with the continuous PC1 credibility measure from Lin and colleagues, published in 2023, which is the first principal component of expert and crowd ratings of news domain credibility on a 0 to 1 scale covering 11,520 domains. Across the 37,020 references that resolved to a score, mean credibility was 0.732, against 0.645 across 159,752 matched first page URLs, a gap of 0.087. The authors note this runs against earlier work finding that AI Overviews drew on lower quality sources, so it is a contested result rather than a settled one.

For anyone buying or selling AI visibility work, the shape of this is uncomfortable in a useful way. A rank tracker reports on the pool the overview partly ignores. Page one is not a prerequisite for the citation above it, and it is not a guarantee of one: a 70.2 percent overlap in one direction still leaves a great many first page results that were never cited. Two systems, two objectives, correlated but not identical, which is the distinction the reference page for Google AI Overviews exists to hold.

First page results

  • 159,752 URLs matched to a credibility score
  • Mean PC1 credibility 0.645
  • 41.4 percent of URLs from user generated platforms
  • Ranked for the query as it was typed

AI Overview references

  • 61,212 URLs from 7,479 hostnames, median 8 per overview
  • Mean PC1 credibility 0.732 across 37,020 scored
  • 14.2 percent of URLs from user generated platforms
  • 29.8 percent of domains on no first page result at all
The two source pools Google selects for the same query, as measured in the paper across the 40 day window.

One claim in nine was not supported by the page it cited

The third measurement is the one to read slowly, because it concerns what the answer says rather than where it points. The authors decomposed each overview into atomic claims, one fact per claim, and checked every claim against the content of the pages that overview cited, using Grok 4.1 Fast Reasoning at temperature 0 and five labels: Clear, Vague, Ambiguous, Incorrect and Omitted.

Across 7,491 verifiable overviews the pipeline returned 98,020 claim level judgments. Of those, 87,204, or 88.97 percent, were consistent: 84.61 percent Clear and 4.36 percent Vague. The remaining 11.03 percent split into Omitted at 6.98 percent, where no cited source mentions the claim at all, Incorrect at 2.66 percent, where a cited source explicitly contradicts the claim it is cited for, and Ambiguous at 1.39 percent, where the cited sources disagree with each other. Omission is roughly 2.6 times more common than contradiction. The median overview had 93.33 percent of its claims grounded, but the tail is real: 205 overviews, 2.74 percent, had fewer than half their claims grounded, and 64, or 0.85 percent, had none at all.

The paper bounds its own error, and that bound should travel with the figure wherever it goes. The pipeline does not extract content from social and video platforms, so a claim supported only by an uncrawled forum or video page scores as unsupported. The authors quantify the exposure: 59.9 percent of inconsistent claims, 5,658 of 9,449, appear in overviews citing at least one such source, and they flag their Climate category number as a measurement artefact of the same gap rather than an AI Overview failure. Publishing the number that weakens your own headline is the part worth copying, and it is the same discipline behind what the evidence actually says about llms.txt, which reports a finding inconvenient for a tool this site ships.

There is a direct consequence for a site that does get cited. A claim can carry your domain as its source while saying something your page does not say, and a reader who sees only the summary has no way to detect that. It is not a crawl problem, and no crawler side check will ever surface it. It is a reason to read the overviews that cite you, which needs no tool at all.

  • Clear 84.61 percent The claim is directly supported by a page the overview cites.
  • Vague 4.36 percent Supported, but loosely. Counted as consistent alongside Clear.
  • Omitted 6.98 percent No cited source mentions the claim at all. The dominant failure mode.
  • Incorrect 2.66 percent A cited source explicitly contradicts the claim it is cited for.
  • Ambiguous 1.39 percent The cited sources disagree with each other.
The five verification labels and their share of 98,020 atomic claims. Xu, Iqbal and Montgomery, 7,491 AI Overviews, 13 March to 21 April 2026.

Who pays for the pages an AI Overview reads

The fourth measurement is economic, and the authors are careful about what it is not. They state that they do not measure traffic loss directly. What they measure is the monetisation of the pages the overviews draw on. Of the 61,212 cited reference URLs, 30,994, or 50.63 percent, displayed visible ads. They call that a conservative lower bound, because a further 14.2 percent of references point at social and video platforms their pipeline does not scrape for advertising markup.

Set against that, the overviews barely displace Google's own inventory. Of the 7,583 results pages carrying an overview, 164, or 2.16 percent, also carried at least one Google sponsored search ad, and 39, or 0.51 percent, placed a sponsored slot above the overview block. No sponsored slot appeared inside the overview container. Where both appear they occupy separate regions of the page, so on this corpus the overview is additive to Google's ad inventory rather than a substitute for it. The authors attribute the low 2.16 percent to a query sample dominated by informational and trending searches rather than commercial ones, which is the right caveat to attach.

One further contrast from the same data set is worth carrying, because it cuts against a common assumption about AI answers. User generated content made up 14.2 percent of AI Overview references, 8,688 of 61,212, against 41.4 percent of first page result URLs, 127,814 of 308,407, on the authors' definition of platforms without strong editorial moderation, which excludes Wikipedia. On this corpus the overview cited proportionally less forum and social content than the results page did, not more.

None of that is a claim about your traffic and this post will not turn it into one. It is a description of the pages that get cited: more often than not, on this sample, they are funded by the page view the answer above them can absorb. If that is your business model, the figure you need comes from your own analytics and from Search Console, not from an external scan, and why no external scan can substitute is an argument this blog has already had to make about a different Google control.

  • Cited pages showing visible ads 50.63% 30,994 of 61,212, stated as a lower bound
  • Cited URLs from social and video platforms 14.2% 8,688 of 61,212
  • First page URLs from the same platforms 41.4% 127,814 of 308,407
  • Overview pages also carrying a sponsored ad 2.16% 164 of 7,583
  • Sponsored slot placed above the overview 0.51% 39 of 7,583
Monetisation measured on the cited pages and on the results pages, from the same 40 day corpus. Xu, Iqbal and Montgomery, arXiv:2605.14021.

What can you actually check on your own site?

Everything above was measured on Google's results pages by somebody else. The honest division of what a scanner adds to it is short, and stating it plainly is more useful than implying the gap is not there.

Start with the eligibility condition Google publishes, because it is first party and it is the shortest path. The AI features documentation says that to be eligible to appear as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, and that there are no additional requirements beyond that. Indexed and snippet eligible is a Googlebot question and a robots meta question, and it is answered inside Search Console rather than by any third party tool, this one included.

What an external scan does answer is whether your text is in the response at all. That is prose parity: the share of your rendered visible text that was already present in the raw HTML. A page that only assembles its content in the browser is a page whose sentences may not exist for any retrieval system that does not run JavaScript, and the failure is silent rather than noisy, which is what makes it worth measuring rather than assuming. What GPTBot sees renders that comparison directly for a single URL, and six real storefront captures show two stores on the same platform producing opposite outcomes on it.

Then check access per crawler rather than as one verdict, because a robots.txt file that admits one fleet and refuses another is the ordinary case rather than an exotic one. The AI crawler checker tests each token separately, the robots.txt tester answers the same question for a specific path and agent, and the bot reference lists which tokens those are. None of them is Googlebot, and that limit is stated here rather than worked around.

After that comes the requirement every answer surface shares, whatever its retrieval looks like: whether a machine can work out what the page is about and who published it. Structured data is what carries that, and our methodology sets out what it is worth in the composite score and why it is worth that rather than more.

Last, the discipline this post has been trying to demonstrate. A grounded model answer is evidence about how models describe your category, not about Google Search, which is the argument in a Gemini citation tool is not an AI Overview tool. A per crawler figure counts requests that claimed to be a crawler, which is the argument in a user agent is a claim, not an identity. Named crawlers with published user agents are the reason the ChatGPT reference page and the Perplexity page can be more concrete than the Google one, and the research page publishes what our own sample supports with the sample size attached. If a vendor tells you it measures your AI Overview presence, ask what it requested and what came back. Anybody who knows their own pipeline can answer that in one sentence.

  • Text in the raw HTML Measurable from outside Prose parity compares rendered visible text against the raw response, per URL.
  • Per crawler robots.txt rules Measurable from outside Each AI crawler token is evaluated separately against the file, for a specific path.
  • Googlebot crawl and index state Not in this scanner Googlebot is not one of the tokens in the registry. Search Console answers this one.
  • Presence in an AI Overview Not observable externally No results page capture and no parser exists here, and no request would reveal it.
What an external scan can and cannot establish about the surface this paper measures.

Related

Common questions

Does ranking on page one get my page into an AI Overview?

Not reliably in either direction. In the Washington University in St. Louis study of 55,393 trending queries run between 13 March and 21 April 2026, 29.8 percent of the domains cited by an AI Overview did not appear anywhere on the first page for the same query, so page one is not a prerequisite. The overlap in the other direction was 70.2 percent, which also means many first page results were not cited. Google's AI features documentation describes a query fan-out technique that issues multiple related searches across subtopics, which is a mechanism consistent with a citation pool that differs from the ranked page.

How often do AI Overviews appear on a search?

On that corpus, 13.7 percent of the time: 7,583 of 55,393 trending queries returned one between 13 March and 21 April 2026. Phrasing mattered more than the average suggests, with question form queries triggering an overview 64.7 percent of the time against 9.5 percent for other phrasings. Category rates ranged from 3.5 percent in Beauty and Fashion to 46.1 percent in Hobbies and Leisure. Those figures describe trending queries in 19 Google Trends categories, not all searches everywhere.

Are the sources an AI Overview cites more or less credible than the normal results?

More credible on the measure that study used, though the finding is contested. Scored with the continuous PC1 domain credibility metric from Lin and colleagues, 2023, AI Overview references averaged 0.732 against 0.645 for matched first page URLs, a gap of 0.087 on a 0 to 1 scale. The authors note this runs against earlier work that found the opposite. Credibility of the source and accuracy of the claim are largely independent in their data, so a credible citation list does not imply a grounded answer.

If an AI Overview misquotes my page, is that something a scanner can detect?

No. In that study 2.66 percent of 98,020 atomic claims were explicitly contradicted by a source the overview itself cited, and a further 6.98 percent were mentioned by no cited source at all. Detecting that requires reading the generated answer next to the page it names, which is a comparison made on Google's output rather than on your server's response. Lantad measures what a crawler receives from your site and makes no claim about what any AI Overview says about it.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.