# A GEO detector flagged 8.90 percent of retrieved pages, and 16.36 percent of 2026 edits

> A benchmark posted to arXiv on 17 August 2026 fetched 10,095 pages out of Google Search and Gemini results between 28 and 31 July 2026, flagged 898 of them as optimised for generative engines, and found that the off-the-shelf detectors it tested were partly reading who wrote a page rather than whether anyone had tuned it.

- Canonical page: https://lantad.co/blog/geo-detector-flagged-8-9-percent-of-retrieved-pages
- This file: https://lantad.co/blog/geo-detector-flagged-8-9-percent-of-retrieved-pages.md
- Last substantive update: 2026-08-20

## Key facts

- **Published:** 2026-08-20
- **Category:** Findings
- **Author:** Lantad
- **Length:** 3610 words
- **Takeaway 1:** GEO-Flag, a paper by Junjie Chu, Ye Leng, Mingjie Li, Yun Shen, Xinyue Shen and Yang Zhang posted to arXiv on 17 August 2026 as arXiv:2608.16824, detected 898 of 10,095 retrieved pages as GEO-optimized, which it reports as an estimated prevalence of 8.90 percent.
- **Takeaway 2:** Those pages were fetched from 28 to 31 July 2026 out of released Google Search and Gemini-grounded retrieval results for 1,000 queries, and the paper reports 8.14 percent prevalence on the Google Search pages against 9.09 percent on the Gemini pages.
- **Takeaway 3:** Among the 1,976 pages exposing a parseable dateModified value, 19.57 percent of the corpus, the paper reports the estimated rate rising from 7.02 percent in 2024 to 12.80 percent in 2025 and 16.36 percent in 2026.
- **Takeaway 4:** The strongest existing detector the paper tested reached an aggregate F1 of 0.880 while scoring 0.375 on its worst subgroup, and the paper attributes that gap to detectors reading whether a page was written by a human or a model rather than whether it was optimised.
- **Takeaway 5:** Lantad ran no part of this study, detects no GEO optimisation of any kind, and measures nothing about which pages an engine retrieves or cites. Every figure below is read from the paper, fetched on 20 August 2026.

## Summary

Almost every argument for making a page machine readable ends somewhere near a citation. [Generative engine optimization](https://lantad.co/glossary/geo) works backwards from that citation instead: it edits the page so an answer engine is more likely to select and quote it. How much of the web that is actually reaching people has been through that treatment is a question people have been guessing at for two years, and the honest answer until now was that nobody had counted it at scale.

A paper posted to arXiv on 17 August 2026 counts it. [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), by Junjie Chu, Ye Leng, Mingjie Li, Yun Shen, Xinyue Shen and Yang Zhang, builds a benchmark, tests whether existing detectors work on it, finds that the best of them are partly reading a different signal from the one they claim to read, trains one that separates the two, and runs the result over pages pulled out of live search results. Lantad ran no part of this study and detects no GEO optimisation of any kind. Everything below is read from the paper, fetched on 20 August 2026, and the one boundary this site can add to it gets its own section rather than being folded quietly into somebody else's numbers.

## What the study counted, and which pages it could count

The paper does two separate things and it is worth keeping them apart, because only one of them produces the 8.90 percent figure. The first is a benchmark called GEOFlagBench: 3,200 webpages spanning 400 queries, four domains named as Health, Finance, Technology and Travel, and eight GEO optimizer families named as GEO Strategy Pool, AutoGEO, PMA, RAID G-SEO, Meta-Optimization, Stealthy GEO Strategy Pool, Stealthy AutoGEO and Human GEO. That corpus is constructed rather than found, which is the point of it: the labels are known because the authors applied the optimisers themselves.

The construction matters for reading everything that follows. For each of the 400 queries the authors built three kinds of control page, giving 1,200 non-GEO controls in total: a human-written page retrieved from Wikipedia or a reliable publisher and dated before October 2022, an AI-polished version of that page, and an AI-generated page built from a structured summary. The polishing step is defined narrowly, as a model that improves grammar, spelling, punctuation and phrasing while preserving the original semantics and content. Applying the eight optimiser families to human-written or AI-generated seeds then yielded 2,000 GEO documents. So the benchmark deliberately contains pages a model touched without optimising them, and that decision turns out to be the whole story of the next section.

The second thing is a deployment on real pages. The authors sampled 1,000 queries by proportional stratified random sampling, using intent labels and query-length bins, from a 5,000-query subset of the ORCAS click dataset, then took the released Google Search and Gemini-grounded retrieval results for those queries and fetched the current page versions from 28 to 31 July 2026. That produced, in the paper's own words, 10,095 usable pages from 13,985 unique URLs, which it reports as 72.18 percent retrieval coverage. Paired analysis across the two channels ran over the 965 queries with at least one usable page from both. That corpus of 10,095 pages, and nothing wider, is what the prevalence figure describes. If you want the field's prior state of knowledge for comparison, the [critical survey of 45 GEO studies](https://lantad.co/blog/geo-survey-45-studies-crawling-stage) posted a month earlier found no stable discoverability gains, and did not attempt a prevalence count at all. The [full paper text](https://arxiv.org/html/2608.16824v1) carries the construction details in more depth than a post can, including the appendix tables this one deliberately does not summarise.

## 8.90 percent overall, and what the 16.36 percent figure actually rests on

The headline is one sentence: the pipeline detects 898 of the 10,095 unique pages as GEO, an estimated prevalence of 8.90 percent. Split by channel, the paper reports GEO detected in 8.14 percent of Google Search pages and 9.09 percent of Gemini pages. That is a real gap between two retrieval systems but a small one, and the paper offers it as a description rather than an explanation. Nothing in it establishes that either system prefers optimised pages, because nothing in it observes how either system ranked or selected them.

The second number is the one most likely to be repeated without its denominator, so it needs stating carefully. The paper reports the estimated rate rising from 7.02 percent in 2024 to 12.80 percent in 2025 and 16.36 percent in 2026. Those three figures are not computed over the 10,095 pages. They are computed over the pages that expose a machine readable modification date, and the paper is explicit about how few of those there are: a parseable dateModified value is available for 1,976 of the 10,095 unique pages, which it gives as 19.57 percent metadata coverage. Four fifths of the corpus is silent about when it last changed and therefore contributes nothing to the trend.

There is a further wrinkle that the paper does not need to state and a reader on this site should. [dateModified](https://schema.org/dateModified) is defined by schema.org as the date on which the CreativeWork was most recently modified, and it is a value the page asserts about itself. Nothing verifies it. A page can carry a 2026 dateModified because a template stamps the build time on every deploy, and a page can carry a 2019 one after a full rewrite. This is the same property that makes [a sitemap lastmod an assertion rather than evidence](https://lantad.co/blog/sitemap-lastmod-is-an-assertion), and it applies with the same force here. The trend is still worth reporting, because a rise from 7.02 to 16.36 across three self-declared cohorts is large and the direction is consistent. It is simply a trend in what pages say about themselves, on the fifth of the corpus that says anything at all, and reading it as a measurement of the web's editing history would be reading past the [structured data](https://lantad.co/glossary/structured-data) rather than through it.

## The best existing detector was partly reading who wrote the page

This is the finding that will outlive the prevalence number, and it is a finding about measurement rather than about the web. The authors ran existing detection methods against GEOFlagBench and the aggregate scores looked respectable. Word TF-IDF with logistic regression achieved the highest F1 of 0.880. Character TF-IDF reached 0.878. Pangram, an AI-text classifier, reached 0.876. A fine-tuned ModernBERT with an 8,192 token window reached 0.862, and zero-shot large language models ranged from 0.770 to 0.776. On an aggregate score alone you would conclude the problem was close to solved.

Then the paper conditions on authorship, and the numbers come apart. Word TF-IDF achieves the highest overall F1 of 0.880, yet its worst-group accuracy is only 0.375, with a false-positive-rate gap of 0.558 between AI-authored and human-authored non-GEO pages. Pangram shows a larger discrepancy still: despite an aggregate F1 of 0.876, its worst-group accuracy is 0.192 and its false-positive-rate gap is 0.767. Read plainly, those detectors were substantially answering the question of whether a model wrote the text, not the question of whether anyone had optimised it for retrieval, and the benchmark exposed it precisely because it contains AI-polished pages that nobody optimised.

The fix the paper proposes is called Intervention-Paired Training, and its mechanism is the interesting part. It supervises the detector on pairs rather than on labels alone. A positive pair requires the GEO score to rise by a margin after an optimiser is applied. A zero pair takes an original page and its AI-polished counterpart and constrains their scores to remain close, which is a direct instruction that polishing is not optimisation. On ModernBERT this lifts F1 from 0.862 to 0.944 and worst-group accuracy from 0.725 to 0.883.

The transferable lesson is uncomfortable for anybody who sells a number, this site included. An aggregate score of 0.876 sitting on top of a worst-subgroup accuracy of 0.192 is not a good measurement with a caveat, it is a measurement that fails on the cases you would most want it for while reporting a healthy average. It is the same failure mode that makes [a withheld grade better than a confident wrong one](https://lantad.co/blog/why-we-withhold-a-grade), and the reason [this scanner publishes how it computes every sub-score](https://lantad.co/methodology) rather than asking anyone to trust the composite. It is also worth separating from a nearby result: this paper is about pages tuned for retrieval, whereas the audit finding that [16 percent of cited sources were AI-generated](https://lantad.co/blog/sixteen-percent-of-cited-sources-were-ai-generated) is about who or what wrote them. GEO-Flag exists partly because those two questions had been getting the same answer from the same tools.

## Optimising a page and being readable are different layers

Here is the part this site can add, and it is a boundary rather than a finding. Every figure in the paper is computed over pages that came back. The corpus is 10,095 usable pages from 13,985 unique URLs, which leaves 3,890 URLs that produced nothing usable after the authors' multi-stage recovery. The paper does not attribute that shortfall to any single cause and neither should anybody reading it: a URL can fail to yield a page because it moved, because it was removed, because it served something that was not an article, or because whatever fetched it did not get through. Whatever the mix, the 8.90 percent describes the readable slice.

That is not a criticism of the study, which states its retrieval coverage plainly in the same sentence as the corpus size. It is a point about what the number can be used for. GEO optimisation is an edit to the words on a page. Whether an [AI crawler](https://lantad.co/glossary/ai-crawler) can obtain those words at all is decided earlier, by the response an unauthenticated fetch receives and by whether the text survives into the initial HTML, which is [the two layer problem this site was built around](https://lantad.co/blog/two-layers-decide-if-ai-can-read-your-site). A page whose body copy only appears after JavaScript runs is not a page with weak optimisation. It is a page with no text to optimise, from the fetcher's point of view, and [prose parity](https://lantad.co/glossary/prose-parity) is the measurement of exactly that gap.

The practical consequence is that these two things sit in a fixed order and only one of them is worth arguing about. Nobody has to decide whether GEO tactics are ethical to find out whether their own pages are retrievable, because retrievability is upstream of the question and answerable in seconds by [asking what a named crawler receives](https://lantad.co/tools/what-gptbot-sees). That is also the question this site keeps a standing sample against: the [crawlability study](https://lantad.co/research/crawlability-study) groups sites by the platform they are built on and scans one page each, so what it offers is whether your kind of site is the readable kind, which is a different question from whether anybody has been tuning their prose. A page that is unreadable cannot be optimised into an answer, and a page that is readable does not need a detector's attention to be cited.

## What the citations on flagged pages looked like

The paper does not stop at flagging. It runs a GEO-gated agent over the pages it detected, auditing the citation URLs those pages carry, on the reasoning that a page written to look well supported is a page worth checking for whether its support exists. Two figures come out of that audit and both are about the flagged pages themselves rather than about search results.

The first: 69.34 percent of citation occurrences on detected GEO pages receive LOW verifiability labels. Split by channel, the LOW share reaches 74.15 percent for Gemini, compared with 45.88 percent for Google Search. The second: the paper reports that 68.84 percent of citation URLs are assigned to C3 sources, where URL Source Tier is described as characterising the publisher behind a citation URL by its publication and review process, and is designed to distinguish sources with strong institutional or editorial control from sources that are easier to create, edit or contribute to. The full C1 to C3 definitions live in an appendix table this post did not read, so the tier labels are reported here without being characterised beyond that stated design intent. Anyone leaning on the C3 figure should open [the paper](https://arxiv.org/abs/2608.16824) and read the table.

The limitation the authors attach to this audit is the one that matters most, and it is stated in their own words: the citation analysis does not assess whether the cited sentence is factually correct, or whether the destination supports the cited claim. So this is a measurement of what a citation points at and whether it resolves, not of whether the citation is honest. That is a narrower claim than the phrase low verifiability suggests on first reading, and the distinction is the same one behind [citation count not being answer influence](https://lantad.co/blog/citation-count-is-not-answer-influence): counting the links in an answer or on a page tells you about the shape of the thing, not about whether it is right. It also sits alongside the earlier finding that [hidden text was the weakest attack on AI search](https://lantad.co/blog/hidden-text-was-the-weakest-attack-on-ai-search), which is a useful corrective to the assumption that manipulation of retrieval is easy in whatever form it takes.

## What this changes for a site owner, and what it does not

The first thing it changes is the vocabulary. Until now the category argument has run on assertion in both directions, with vendors claiming that optimising for answer engines is a discipline and critics claiming it is spam, and neither side able to say how much of it is out there. There is now a published estimate with a method attached, a date on it, and an author list. Whether 8.90 percent is high or low is a judgement, and reasonable people will read the same number in opposite directions.

The second thing it changes is what a detector claim is worth. If you are evaluating any tool that says it can tell you whether a page has been optimised, or whether a competitor is doing it, the question this paper hands you is not what its accuracy is. It is what its accuracy is on the subgroup you care about, and specifically whether it can tell an optimised page from a page a model merely tidied up. On the evidence here, two of the strongest general-purpose methods could not, and their aggregate scores gave no hint of it.

What it does not change is anything about how a page gets read in the first place. None of these figures describe [answer engine optimisation](https://lantad.co/glossary/aeo) as a practice that works, and the paper makes no claim that a flagged page was actually cited more often. Google's own position on this class of tactic has not moved either, and the [five GEO tactics Google names and tells you to ignore](https://lantad.co/blog/google-names-five-geo-tactics-to-ignore) are still the published guidance. If the reason you are reading this is that you want to be present in answers, the sequence that actually applies starts with the ordinary questions: which [AI crawler tokens](https://lantad.co/tools/ai-crawlers) your robots.txt names, whether your text is in the HTML those crawlers receive, and then what the platform documentation asks for, which for the largest surface is set out in [how to get cited in Google AI Overviews](https://lantad.co/how-to-get-cited/google-ai-overviews).

The last thing worth saying is what the paper says about itself. It states that its detector's performance against unseen GEO strategies is uncertain, that its real-world estimates have no ground truth so detector errors propagate into the prevalence figure, and that URLs change over time as pages are edited, removed or archived. Those are the right limitations to publish and they are all the ones a critical reader would have raised. They also mean the 8.90 percent is a first estimate rather than a settled figure, and anyone quoting it in six months should say which paper and which July it came from.

## Questions and answers

**How much of the web is GEO-optimized?**

The only published estimate is from GEO-Flag, arXiv:2608.16824, posted 17 August 2026, which detected 898 of 10,095 pages as GEO for an estimated prevalence of 8.90 percent. Those pages were fetched from 28 to 31 July 2026 out of released Google Search and Gemini-grounded retrieval results for 1,000 queries, so the figure describes retrieved pages rather than the web as a whole, and the authors state it has no ground truth.

**Why is the 2026 figure of 16.36 percent so much higher than 8.90 percent?**

Because it has a different denominator. The year figures are computed only over pages exposing a parseable dateModified value, which the paper reports as 1,976 of the 10,095 pages, or 19.57 percent metadata coverage. The rate reported over that subset rises from 7.02 percent in 2024 to 12.80 percent in 2025 and 16.36 percent in 2026. dateModified is also a value a page asserts about itself and nothing verifies it.

**Can an AI-text detector tell whether a page has been optimised for answer engines?**

Not reliably, on the evidence in this paper. Pangram reached an aggregate F1 of 0.876 on the benchmark while scoring 0.192 on its worst authorship-conditioned subgroup, with a false-positive-rate gap of 0.767 between AI-authored and human-authored pages that nobody had optimised. The authors read that as detectors answering whether a model wrote the text rather than whether it was optimised, and trained a method specifically to separate the two.

**Did Lantad measure any of this?**

No. Lantad detects no GEO optimisation, scores no page for it, and measures nothing about which pages a search or answer engine retrieves or cites. Every figure in this post is read from arXiv:2608.16824, fetched on 20 August 2026. What Lantad measures is upstream of all of it: what a named AI crawler receives when it requests your URL, and how much of the rendered text survives into that response.

---

Lantad measures whether AI crawlers can actually read a page: it fetches as a non-rendering
crawler, renders as a browser, and reports the gap. Free scan, one URL, no signup.

Method and weights: https://lantad.co/methodology | All pages as markdown: https://lantad.co/md | Crawler policy: https://lantad.co/bot
