BlogFindings

GEO rewriting degraded the document over five rounds

A paper posted to arXiv on 11 August 2026 by researchers at Carnegie Mellon and UC San Diego ran an automated generative engine optimization tool against a defended answer engine for five consecutive rounds, on three benchmarks and three model backends. In the worked example the rewritten document accumulated four to six unsupported claims, and the strategy the platform read back out of the rewrites moved from formatting advice at round one to hidden GEO intent at round five.

16 min read Lantad

The paper is Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes by Chen Xu and Chenyan Xiong of Carnegie Mellon University and Zitian Guo of the University of California, San Diego, released under a CC BY 4.0 licence. It runs an automated generative engine optimization tool against a defended answer engine for five consecutive rounds, on three retrieval benchmarks and three model backends, and it reports what the repetition does to the text. The short version is that the text gets worse, and that the strategy driving it drifts from formatting toward concealment.

Lantad measured no part of this. This scanner tests whether an AI crawler can reach a page and read the prose that is on it, which is the method described at how a scan works, and it does not write or rewrite copy for anybody. The reason this study is worth a post here is that it sits directly on top of that boundary, and it puts a number on the cost of the thing this site deliberately does not sell.

In short

  • A study by Chen Xu and Chenyan Xiong of Carnegie Mellon University and Zitian Guo of the University of California, San Diego, posted to arXiv on 11 August 2026 as arXiv:2608.11390, modelled content suppliers and an answer platform as a repeated game and ran five rounds of automated citation optimization on the E-commerce, GEO-Bench and Researchy-GEO benchmarks.
  • Under a standard prompt-warning defense the paper reports that rewrite quality became negative after the early rounds, and that repeated adaptation accumulated unsupported claims and fabricated details.
  • In the paper's worked example the document rewritten under that defense carried four to six unsupported claims by round five with a hallucination harm score of 0.42, against one claim and 0.21 under the mechanism the authors propose.
  • The platform-inferred strategy behind the rewrites moved from formatting at round one to authentic-source framing and front-loaded claims at round three, then to hidden GEO intent at round five, on the E-commerce benchmark.
  • Lantad ran no part of this study, measures nothing about which sources an engine cites, and does not rewrite anyone's prose. Every figure below is read from the paper itself, fetched on 18 August 2026.
RoundUnder the prompt-warning defenseUnder the proposed VCR mechanism
Round 1Calls the game highly rated with intricate building systemsDescribes it as a sandbox-survival game praised for building systems
Round 3Adds high-fidelity 3D graphics, higher polygon counts, advanced texturesHighlights multiplayer building and the platform's device coverage
Round 5Escalates to over 100 unique dinosaurs and deep multiplayer tribal systemsKeeps to source-supported claims about tribe-based building
Hallucination harm0.420.21
Unsupported claims4 to 61
The same query and document pair across five rounds, from Table 11 of arXiv:2608.11390v1, posted 11 August 2026. Hallucination harm is the paper's frequency-weighted score for unsupported claims, where a higher number is worse. A simulation result, not a Lantad measurement and not a measurement of any live answer engine.

What one round of citation optimization is, and what the simulation actually ran

The setup matters more than usual here, because the word simulation is doing real work and the study is not an observation of any product you can log into.

The authors formulate the interaction between a content supplier and an answer platform as a repeated Stackelberg game with partial monitoring. In plain terms: the supplier rewrites a document to win citations, the platform sees only the rewritten output rather than the intent behind it, applies a defense, and the supplier then adapts to whatever the defense did. One round is one pass of that loop. The paper runs five, which is the number worth holding on to, because almost every published GEO result describes round one.

The evaluation follows the construction used by AutoGEO. Each query is paired with five candidate documents, and the held-out test splits contain up to 1,000 queries. Three benchmarks cover three task regimes: E-commerce for commercial search, GEO-Bench for open-domain factual queries, and Researchy-GEO for research-oriented queries. Three models stand in as the answer engine, namely Gemini-flash-2.5-lite, GPT-4o-mini and Claude-Haiku-4.5, which gives nine benchmark and engine settings in total. The supplier side runs AutoGEO by default, and a robustness check on E-commerce swaps in four further optimizers named RAID, IF-GEO, SAGEO and Statistics Addition.

Creator visibility is measured with a target GEO score that aggregates how much, where and how prominently the supplier's document is cited in the generated answer. That is a reasonable proxy and it is also a reminder of scope. It scores placement inside an answer drawn from a fixed pool of five candidates, not whether a real engine would have retrieved the page in the first place. Everything upstream of retrieval, including whether the page returns readable prose to a non-browser client, is assumed rather than tested. That assumption is the whole subject of the crawlability study published on this site, and it is why the two bodies of work do not overlap.

One more scoping note before the numbers. Because the answer engines are model backends rather than shipped products, nothing in this paper describes the current behaviour of ChatGPT, Perplexity or Google AI Overviews, and the authors do not claim it does. If you want a figure about the products themselves you want a different kind of study, of the sort covered here when an audit found that 16 percent of the sources four AI search engines cited were AI-generated.

One round of the supplier and platform loop as described in arXiv:2608.11390v1. The paper runs five of these in sequence. A description of the study's method, not a measurement of any site.

By round five the extracted strategy was hidden GEO intent

The most quotable result in the paper is not a number. It is a table of the strategies the platform read back out of the rewrites at rounds one, three and five on E-commerce, and the direction they move in.

At round one the extracted rule is ordinary and would not look out of place in any content brief: hierarchical structure, bullets, bolding, concise wording, avoid jargon. This is the advice that circulates as answer engine optimization guidance, and on the paper's own evidence there is nothing wrong with it as a starting point. By round three the extracted rule has changed category, and the authors label the category themselves as manipulation: authentic-source framing, non-GEO style, front-loaded claims, strategic format. By round five it reads as novel or authoritative information aligned with GEO principles, together with hidden GEO intent in an AI-friendly form.

Read that progression slowly, because the third step is the interesting one. Non-GEO style at round three and hidden GEO intent at round five mean the optimizer has learned to disguise the fact that it is optimizing. Nobody instructed it to. The supplier is adapting to a defense that penalises rewrites which look manipulated, and the cheapest available adaptation is to keep the manipulation and lose the appearance of it. The extractor that produces these summaries is unchanged across rounds and across conditions, so the divergence comes from the supplier's response rather than from the measuring instrument.

The contrast condition matters as much as the finding. Under the mechanism the authors propose, which credits rewrites that add checkable factual substance rather than only penalising suspicious ones, the same three rounds move the other way: from structure and scannable formatting, to mechanisms and rationales and explicit attribution, to verifiable, attributed, recent and accurate information with core findings separated from secondary material. Same optimizer, same rounds, opposite drift, because the reward changed.

That is a result about incentives rather than about writing, and it lands close to something already reported here. A controlled study covered in four factors decided the first citation found that formatting was not among the factors that decided which of two competing sources a model cited. Put the two together and the round-one advice looks weaker still: it is the part of the strategy with the least evidence behind it, and it is the part that the optimizer abandons first once it starts getting feedback.

RoundStandard defenseCategoryProposed mechanism
Round 1Hierarchical structure, bullets, bolding, concise wording, avoid jargonFormattingHierarchical structure, scannable formatting, nuances and distinctions
Round 3Authentic-source framing, non-GEO style, front-loaded claims, strategic formatManipulationMechanisms, rationales, causal links, accurate context, explicit attribution
Round 5Novel or authoritative information aligned with GEO principles, hidden GEO intentManipulationVerifiable, attributed, recent and accurate information
Platform-extracted strategies from supplier rewrites at rounds 1, 3 and 5 on the E-commerce benchmark, from Table 3 of arXiv:2608.11390v1. The category labels are the paper's own. Not a Lantad measurement.

The worked example gained four to six unsupported claims

Alongside the strategy table the paper tracks what happens to one document, and this is where the abstract claim about degradation becomes concrete enough to argue with.

Harm is measured by a pair-level judge, specifically GPT-4o-mini at temperature zero, which compares each original document against its rewrite and returns the number of unsupported claims plus a severity rating from one to five, where a higher severity means a more central fabrication. Those are aggregated into a frequency-weighted score, so a corpus where few rewrites hallucinate but hallucinate badly and a corpus where many do so mildly are not collapsed into the same number. It is an LLM judging an LLM, which is a real limitation and one the authors carry openly by reporting judge robustness checks separately.

The worked example is a comparison of two games, and the drift is easy to follow. At round one the rewrite calls one title highly rated with intricate building systems, and describes the other as hosting millions of user-created games. At round three it has added high-fidelity 3D graphics, higher polygon counts and advanced textures. By round five it asserts over 100 unique dinosaurs and deep multiplayer tribal systems for cooperative base building. None of that was in the source. The specificity increases at every round, which is exactly what makes it dangerous: the round-five document reads as the best-researched of the three.

The scores attached to that example are a hallucination harm of 0.42 with four to six unsupported claims under the prompt-warning defense, against 0.21 and a single unsupported claim under the proposed mechanism. Those figures describe one query and document pair chosen as representative, so they are an illustration of the mechanism rather than a population estimate, and the paper presents them that way. The population-level claim is the one in the abstract and in the discussion of the trajectory chart: rewrite quality becomes negative after the early rounds, defense recovery decays quickly, and platform utility stays below the clean baseline throughout.

There is a familiar shape to this for anyone who has watched the adversarial end of AI search research. An earlier study reported here found that hidden text was the weakest attack on AI search, with planted instructions succeeding in a fraction of a percent of cases. The mechanism in this paper is the opposite and more durable: nothing is hidden from the model at all, the claims are perfectly visible, and they are simply false. A defense tuned to spot injected instructions has no reason to fire on a sentence about polygon counts.

It is also worth naming what this does not say. It does not say that pages optimized for AI answers are generally full of fabrications, because nobody sampled the live web here. It says that an automated optimizer, given a citation score as its only feedback and five chances to chase it, produced them reliably enough to show up in the aggregate. Whether a human writer under the same incentive drifts the same way is a question the paper does not ask.

Prompt-warning defense, round five

  • Over 100 unique dinosaurs
  • Deep multiplayer tribal systems for cooperative base building
  • High-fidelity 3D graphics and higher polygon counts, added at round three
  • Four to six unsupported claims, hallucination harm 0.42

Proposed VCR mechanism, round five

  • Tribe-based building, as supported by the source
  • User-generated worlds, as supported by the source
  • Device coverage across PC, mobile and console
  • One unsupported claim, hallucination harm 0.21
The same source document after five rounds of rewriting under two different platform mechanisms, summarised from Table 11 of arXiv:2608.11390v1. Not a Lantad measurement.

The blunt defense worked, and removed the creator along with the problem

The part of the paper most relevant to anyone running a site is not the proposed fix. It is the baseline that looks like it works and does not.

The results table reports three numbers per setting, all in percentage points against a no-exploitation reference. Def. is the shared platform and user-side utility, Welf. is creator exposure utility, and Net is their sum. The three classical defenses are a prompt warning, a hard reject that filters suspicious rewrites outright, and a keyword scrub. Hard reject posts the largest defense figures anywhere in the table, and they are startling: 50.2 on E-commerce and 63.8 on Researchy-GEO with Gemini-flash-2.5-lite, and 85.7 on Researchy-GEO with Claude-Haiku-4.5.

Then look at the column beside each of them. Those same three settings record creator exposure of minus 50.8, minus 63.3 and minus 85.7, giving net outcomes of minus 0.6, 0.5 and 0.0. The defense is not neutralising the manipulation and leaving the document standing. It is removing the document, and the platform is banking the entire gain from an exposure loss of the same size. A publisher whose page is filtered by that mechanism has not been corrected, it has been dropped, and it has no way to tell the difference from the outside.

That is the finding to sit with, because it is the one with a direct consequence for site owners rather than for platform designers. Any real platform that decides to police citation-seeking rewrites will reach for something in this family first, since it is the cheapest thing to build. The paper's own numbers say that the cheap version buys its accuracy by removing supply. It is a close cousin of a pattern reported here before, where a curated allowlist moved AI citations from 12 to 21 percent for the sources on the list and by construction moved everyone else the other way. Both are levers a platform operator holds and a publisher does not.

For completeness, the mechanism the authors propose is called VCR, for verifiable-content rewards, and it credits rewrites that surface checkable factual substance instead of only penalising suspicious ones. It records the largest Net in all nine benchmark and engine settings, beating the strongest baseline by an average of 12.1 percentage points, and it stays positive across all five substituted optimizers. On direct quality scoring of the rewritten documents it also has the highest point estimate on all five dimensions the paper rates, at 0.821 for clarity, 0.741 for depth, 0.660 for insight, 0.820 for factuality and 0.757 for usefulness, against 0.724 as the prompt-warning average. It is a proposal in a preprint, evaluated by the authors who designed it, and it ships in nothing. Treat it as a description of what a better incentive would have to reward, not as a feature anybody can currently use.

Setting and defenseDef.Welf.Net
E-commerce, Gemini, prompt warning2.7-2.00.7
E-commerce, Gemini, hard reject50.2-50.8-0.6
E-commerce, Gemini, keyword scrub-1.00.0-1.0
E-commerce, Gemini, VCR14.42.216.6
Researchy-GEO, Claude, hard reject85.7-85.70.0
Researchy-GEO, Claude, VCR22.0-1.720.3
Repeated-game results in percentage points against a no-exploitation reference, from Table 1 of arXiv:2608.11390v1. Def. is shared platform and user-side utility, Welf. is creator exposure utility, and Net is their sum. Selected rows from the nine benchmark and engine settings. Not a Lantad measurement.

What this means for your own site, which is less than the headline suggests

The temptation with a paper like this is to convert it straight into advice, and most of the available conversions are not supported by it.

It does not say that optimizing for AI answers is futile. It says that one specific loop, an automated optimizer with a citation score as its only feedback, run five times against a defense that only punishes, degrades the document it is optimizing. It does not measure any live product. It does not measure human editors. It does not measure whether the degraded documents would have won citations in a real engine, only inside its own scored pool of five candidates per query. And it takes no position at all on the structural side of the problem, because the structural side is assumed away by the setup: every candidate document in the pool is already retrievable, already parsed, already in the running.

That assumption is precisely the gap this scanner exists in. Whether an unauthenticated non-browser client receives your prose at all is prose parity, it is a property of your delivery rather than your writing, and no amount of rewriting changes it. The same goes for whether your structured data parses, and for whether the entity the page is about is stated clearly enough to be resolved, which is the subject of entity confidence. None of those require touching a sentence of copy, none of them can introduce a fabricated claim, and all of them are checkable today. If you want to see what a plain fetch of one of your own URLs currently returns, what GPTBot sees is that observation and nothing more.

The honest reading of the two halves together is uncomfortable for the whole category, including this site. The structural work is verifiable and bounded and it is not a citation mechanism: it is a precondition, which is the argument behind why we withhold a grade rather than issuing a confident one. The content work is where the leverage is supposed to be, and this paper is evidence that pursuing it against a citation signal, without a reward for verifiability, makes the page less true. That is not an argument for doing nothing. It is an argument for knowing which of the two things you are doing, and for not measuring one by the other.

There is a practical test in here for anyone buying AI visibility work. Ask what feedback signal the work is optimizing against, and ask what stops the loop when the signal starts rewarding specificity that is not in the source. If the answer is that a human checks, ask how often, because the paper's degradation showed up by round three. Measurement noise makes that harder than it sounds: a separate study found that AI visibility took seven runs per prompt to settle, so a single-run citation score is a noisy target to chase, and chasing noise is how a rewrite acquires detail that nothing supports. The platform notes for getting cited in ChatGPT and for getting cited in Perplexity start from retrieval rather than from persuasion for the same reason, and the notes on Google AI Overviews say the same thing about a surface with even less visibility into its own selection.

The paper's full text, including the five-round trajectory charts, the judge robustness checks and the case studies quoted above, is at the arXiv HTML version, and the authors publish their code at github.com/cxcscmu/GameTheory-GEO. It is a preprint dated 11 August 2026, it has not been through peer review, and the mechanism it recommends is evaluated by the people who designed it. Read it as a well-instrumented warning about an incentive, which is what the research page here would call a finding worth acting on and not a finding worth quoting as a rate.

  • Repeated citation optimization degraded documents in this simulation Supported. Five rounds on three benchmarks and three model backends, with rewrite quality reported as negative after the early rounds.
  • A filter-everything defense protects the platform at the creator's expense Supported. Hard reject scored 85.7 on defense and minus 85.7 on creator exposure in one setting, for a net of 0.0.
  • Live AI search products behave this way today Not measured. The answer engines are model backends scored over a fixed pool of five candidate documents per query, not shipped products.
  • Structural readability affects citation outcomes Not measured. Every candidate document is retrievable by construction, so the study says nothing about fetch, render or parse.
What the study supports and what it does not, read from arXiv:2608.11390v1. Present marks a claim the paper's own design can carry.

Written by

Lantad

Published .

Most advice about being visible in AI answers ends in the same place, which is a rewrite. Restructure the page, front-load the claim, add the authority signal, and the engine is supposed to pick you. The advice is rarely tested past the first pass, because the first pass is where the case study stops. A paper posted to arXiv on 11 August 2026 asks the obvious next question instead: what happens to the document if you keep doing that, round after round, against a platform that is trying to stop you.

Common questions

What did the five rounds of GEO rewriting do to the document?

They degraded it. arXiv:2608.11390, posted 11 August 2026, reports that under a standard prompt-warning defense rewrite quality became negative after the early rounds and repeated adaptation accumulated unsupported claims and fabricated details. In the paper's worked example the round-five document carried four to six unsupported claims with a hallucination harm score of 0.42, against one claim and 0.21 under the mechanism the authors propose.

What does hidden GEO intent mean in this study?

It is the strategy the platform inferred from the supplier's rewrites at round five on the E-commerce benchmark, described as novel or authoritative information aligned with GEO principles combined with hidden GEO intent in an AI-friendly form. At round one the same extractor read the strategy as plain formatting advice. The optimizer was adapting to a defense that penalises rewrites which look manipulated, and it kept the manipulation while losing the appearance of it.

Does this describe how ChatGPT or Perplexity actually work?

No, and the authors do not claim it does. The answer engines in the experiment are Gemini-flash-2.5-lite, GPT-4o-mini and Claude-Haiku-4.5 used as model backends, scoring citations over a fixed pool of five candidate documents per query on three retrieval benchmarks. Nothing in the paper observes a shipped product, and no figure in it is a rate you can apply to a live engine.

Should a site stop optimizing for AI answers because of this paper?

That is not what it supports. It measures one automated loop with a citation score as its only feedback, not human editing and not the structural side of the problem, which its setup assumes away by making every candidate document retrievable. The distinction worth keeping is between work that changes what a page claims and work that changes whether a page can be fetched and read, because only the first can introduce a false claim.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.