# Generative engine optimization: structure moved citations 7.8 points, not 17.3

> A paper posted to arXiv on 31 March 2026 reports that editing document structure alone raised citation rate across six generative engines. Its own table puts the move at 45.0 percent to 52.8 percent, and the ablation attributes 84.6 percent of that gain to heading hierarchy and chunking rather than to emphasis.

- Canonical page: https://lantad.co/blog/generative-engine-optimization-structure-7-8-points
- This file: https://lantad.co/blog/generative-engine-optimization-structure-7-8-points.md
- Last substantive update: 2026-09-02

## Key facts

- **Published:** 2026-09-02
- **Category:** Findings
- **Author:** Lantad
- **Length:** 3109 words
- **Takeaway 1:** A paper posted to arXiv on 31 March 2026 as arXiv:2603.29979, by four authors at the University of Tokyo, the University of Tsukuba, Hiroshima University and the National Institute of Informatics, reports that editing document structure alone raised citation rate from 45.0 percent to 52.8 percent across six generative engines.
- **Takeaway 2:** That distance is 7.8 percentage points, and the paper's headline 17.3 percent is the same distance expressed against the 45.0 percent baseline, which its own ablation table records as a drop of 7.8 from the optimised figure.
- **Takeaway 3:** The paper's generative engine optimization ablation attributes 44.9 percent of the gain to macro-structure, defined as heading hierarchy, navigational elements and document flow, 39.7 percent to meso-structure, being paragraph organisation, lists and tables, and 15.4 percent to micro-structure, being emphasis markers, keyword placement and syntactic patterns.
- **Takeaway 4:** The result rests on 200 articles sampled from GEO-bench across six domains, paired with 377 queries and run against six named platforms in two versions each for 2,400 test cases, with a paired t-test at p below 0.001 and a reported Cohen's d of 0.64.
- **Takeaway 5:** Lantad ran none of these trials and holds no citation data. The paper does not state when its experiments ran, reports no per-platform results, and carries no limitations section, so what follows is a reading of it rather than a measurement by this scanner.

## Summary

Most [generative engine optimization](https://lantad.co/glossary/geo) advice is advice about words: add a statistic, quote a named source, answer the question in the reader's own phrasing. A paper posted to arXiv at the end of March takes the opposite cut. It holds the words still, changes only the shape of the document, and asks whether an answer engine cites it more often. The answer it reports is yes, and the number it leads with is 17.3 percent.

That number is worth opening rather than repeating, because it does not mean what a quick reading suggests. The paper's own results table records citation rate moving from 45.0 percent to 52.8 percent, a distance of 7.8 percentage points, and 17.3 percent is that distance expressed against the 45.0 baseline. Both figures are in the paper and only one of them is in the abstract. What follows reports what the experiment did, what its ablation attributes to which layer of structure, and why it sits alongside rather than against [the controlled study read here in August](https://lantad.co/blog/four-factors-decided-the-first-citation) that found formatting did not move the first citation at all.

## What did this generative engine optimization experiment actually measure?

[Structural Feature Engineering for Generative Engine Optimization](https://arxiv.org/abs/2603.29979) was posted to arXiv on 31 March 2026 by Junwei Yu of the University of Tokyo, Yang MuFeng of the University of Tsukuba, Yepeng Ding of Hiroshima University and Hiroyuki Sato of the University of Tokyo and the National Institute of Informatics. It is filed under computation and language, human-computer interaction and information retrieval, it carries a CC BY-NC-SA 4.0 licence, and the version on arXiv is v1 with no revision.

The dataset is small and the paper says so plainly. Two hundred articles were sampled, approximately 33 per domain, from GEO-bench, across six domains named as Biography, Health, Technology, Finance, Travel and Science. Those articles are paired with 377 real-world queries and have a mean length of 2,547 words. Each article exists in two versions, the original and a structurally edited one, and each version was put to six platforms, which is where the figure of 2,400 test cases comes from.

The step that makes the result interesting is the one that holds meaning constant. The authors state that semantic preservation was verified with Bge-m3 sentence embeddings at a mean similarity of 0.843 between the original and the edited version. That is a check on the edits, not a proof that nothing semantic changed: an embedding similarity of 0.843 leaves room for real differences in wording. It is worth stating what it does establish, which is that the intervention was not a rewrite in disguise, and worth not stating more.

Two metrics carry the results. Citation rate is defined in the paper as the number of queries citing the content divided by total queries, which is a count of outcomes and needs no interpretation. Visibility Score is defined as 0.4 times Coverage plus 0.3 times Position plus 0.3 times Influence. Those three weights are a choice the authors made, in the same way that the weights behind [our own composite score](https://lantad.co/methodology) are a choice rather than a finding, so a movement in Visibility Score is a movement in a formula somebody designed and citation rate is the harder number of the two.

The six platforms are named, and they are grouped rather than reported individually, which matters later.

## Is a 17.3 percent lift a gain of 17.3 points?

It is not, and the paper contains everything needed to see that.

Table IV gives an overall baseline citation rate of 45.0 percent and an optimised citation rate of 52.8 percent. The difference between those two is 7.8 percentage points. Divide 7.8 by the 45.0 baseline and the result is 0.1733, which is the 17.3 percent in the abstract. The paper is not doing anything unusual here: reporting a relative improvement is ordinary practice in this literature. The trap is that the sentence "improvements in citation rate (17.3 percent)" reads, to somebody scanning an abstract, like a document that was cited 45 percent of the time now being cited about 62 percent of the time.

The paper's own ablation table settles the arithmetic without needing any of ours. Its baseline row carries a Drop column reading minus 7.8 against the fully optimised figure of 52.8. The same 7.8 points appears in both tables, once as an absolute distance and once, in the abstract, expressed relative to where it started.

The per-group figures behave the same way. Search-then-Synthesize moves 43.7 to 52.1, which is 8.4 points and the paper's +19.2 percent. Iterative Refinement moves 52.3 to 59.6, which is 7.3 points and +14.0 percent. Integrated Search-Generation moves 39.1 to 46.8, which is 7.7 points and +19.7 percent. Notice that the largest relative gain belongs to the group with the lowest baseline, which is what relative gains do, and that the three absolute movements sit within about a point of each other. A reader told only the percentages would conclude that Integrated Search-Generation responded to structure roughly 40 percent more strongly than Iterative Refinement. A reader given the points would conclude the three responded to structure by about the same amount.

The statistics reported alongside are a paired t-test at p below 0.001 with n of 200 and a Cohen's d of 0.64, which the paper describes as a medium-to-large effect size. None of that is in dispute here. The point is narrower and it is a point about reading: an effect can be real, significant and worth acting on while being half the size a headline implies. This is the same care behind reporting that [title and meta rewrites gained 22 percent of retrievals in a GEO benchmark](https://lantad.co/blog/metadata-rewrites-gained-22-percent-of-retrievals) in the benchmark's own terms, and it comes from the same position as [refusing to print a grade for a page that could not be measured](https://lantad.co/blog/why-we-withhold-a-grade): a number is only as useful as the shape it was measured in.

## Which level of structure carried the gain?

The paper decomposes structure into three levels and then removes them one at a time, which is the most useful thing in it for anyone deciding what to change first.

Macro-structure is document-level architecture: heading hierarchy, navigational elements and document flow. Meso-structure is section-level organisation: paragraph organisation, list structures and table formatting. Micro-structure is sentence-level: emphasis markers, keyword placement and syntactic patterns. Removing macro-structure from the optimisation drops citation rate from 52.8 to 49.3, a loss of 3.5 points. Removing meso-structure drops it to 49.7, a loss of 3.1. Removing micro-structure drops it to 51.6, a loss of 1.2. The paper turns those into contribution shares of 44.9, 39.7 and 15.4 percent.

One caveat belongs on that decomposition before anybody quotes it. Those three shares sum to exactly 100.0 percent, and the paper reads that as evidence the levels operate largely independently. An ablation that removes one level at a time and then normalises the losses will produce shares summing to 100 whether the levels interact or not, so the sum is a property of the arithmetic rather than a test of independence. The ordering is the finding. The tidiness is not.

Read as an ordering, it says something a working site owner can use: the tier that people reach for first, which is bolding, keyword placement and sentence shape, is the tier the paper attributes least to, at 1.2 points of the 7.8. Heading hierarchy and document flow are worth almost three times that.

Those are the things this scanner already looks at, and what it has found on real pages is not encouraging. [An h1 inside a header element counted as no h1 at all](https://lantad.co/blog/an-h1-inside-a-header-counted-as-none) in the extractor, which is a heading hierarchy that exists visually and not structurally. [Deleting the main element cost 1,665 of 13,615 words](https://lantad.co/blog/deleting-main-cost-1665-of-13615-words) across the five real pages captured on 15 July 2026, which is document flow measured as a quantity of text. On the same five pages, [1,395 of 2,729 text blocks were navigation rather than main content](https://lantad.co/blog/half-the-text-blocks-were-navigation), and [all three table elements present were navigation](https://lantad.co/blog/three-tables-on-five-pages-all-navigation) rather than data. [Thirty-one disclosure widgets held no FAQ](https://lantad.co/blog/thirty-one-disclosure-widgets-held-no-faq) between them, and [55 of 245 headings carried an id](https://lantad.co/blog/read-more-deep-links-and-heading-ids) a deep link could target. Every one of those is a macro or meso property in the paper's vocabulary, and none of them is [structured data](https://lantad.co/glossary/structured-data) in the schema sense, which is a separate layer with its own separate problems.

## Why this does not contradict the study where formatting did not matter

Three weeks ago this blog reported a study that appears to say the opposite. [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), posted to arXiv on 25 May 2026, ran 252,000 trials across six language models, showing each model two competing sources differing in exactly one of 18 factors. For structured against dense formatting it reported odds ratios of 1.68, 1.03, 0.79, 0.90, 0.78 and 1.25 across the six models, against its own scale that calls anything below 1.5 negligible. Formatting was among the factors with no consistent effect.

Set side by side, one paper says structure is worth 7.8 points of citation rate and the other says formatting does not move the first citation. Both can be true, and the reason is in the experimental designs rather than in either result.

Every trial in the May study handed the model the full text of both candidate sources through a simulated tool response, and the paper states that no search engine was ever called. The document was already in the context window. What that experiment measured is selection: given two documents a model can both see completely, which does it cite first.

The March study ran against six live commercial platforms that performed their own retrieval. Its citation rate counts the share of queries in which the document was cited at all, which folds retrieval, chunking and selection into a single outcome. Structure therefore has somewhere to act in the second design that it does not have in the first, which is everything that happens before the document reaches the model: what a retriever indexes, where a chunker cuts, which passage a reranker surfaces.

That reading makes both results possible at once. It is a reading and not a finding, and neither paper demonstrates it, so it should not be repeated as though one of them did. What can be said without inventing anything is narrower: the two experiments measured different stages, and a claim about one is not a claim about the other. This is the same gap [the survey of 45 GEO studies](https://lantad.co/blog/geo-survey-45-studies-crawling-stage) identified when it put crawling second in its own pipeline and reported that few of the studies it reviewed observe that stage at all. It is also why the measurement this scanner does perform sits earlier still, at [whether the prose a browser renders is the prose a crawler receives](https://lantad.co/glossary/prose-parity), which is a question that has to resolve before either paper's subject begins. You can see that comparison for a page directly with [the crawler view tool](https://lantad.co/tools/what-gptbot-sees).

## What the paper does not report, and what to check on your own site

Four gaps are worth naming, because they decide how far the result travels.

It carries no date for the experiments. A citation rate measured against six commercial products is a measurement of a moment, and the paper does not name the moment. That is not a small omission for this subject: the products in its platform table ship changes continuously, and a rate observed in one month is not a rate that holds in another. Dating every figure is [the standing rule for research published here](https://lantad.co/research) for exactly this reason.

It reports no per-platform results. Table IV aggregates the six platforms into three architectural groups, so nothing in the paper says what ChatGPT did, or Perplexity, or any single engine. The architecture grouping is the authors' own classification, and a reader who wants to know whether one platform rewards headings more than another will not find it here.

It carries no limitations or threats-to-validity section. The paper runs from its introduction through related work, the framework, the experimental evaluation and its ablation, straight to a conclusion. Searching the text for either word returns nothing. That is worth stating flatly rather than as an accusation, because the caveats above are ones a limitations section would normally have raised.

And the second headline number is scored by a language model. The 18.5 percent improvement in subjective quality comes from G-Eval using Gemini 2.5 Pro across seven dimensions, taking the median of five samples at temperature 0.3. Its two largest components are Influence at plus 32.0 percent and Click Probability at plus 31.4 percent. A model's estimate of click probability is not a click, and reporting it beside a counted citation rate invites a reader to weigh the two equally when only one of them counted anything.

Two more limits belong to the material rather than the write-up. The articles average 2,547 words, so this is evidence about long-form documents and not about a product page, a pricing page or a docs page. And single-run citation rates against live engines are noisy in a way this design does not address: measuring the same prompt repeatedly, [AI visibility took seven runs per prompt to settle](https://lantad.co/blog/ai-visibility-took-seven-runs-per-prompt-to-settle) in the earlier work read here.

What survives all of that is still useful, and it is the ordering rather than the size. If the structural tiers really do rank macro, then meso, then micro, the cheapest work on any page is the work furthest from the words: one h1 that the parser can find, headings that describe sections rather than decorate them, a main element that contains the article, tables that hold data rather than links. Those are checkable in a browser's element inspector in a few minutes, and they are checkable from outside because they are in the HTML that [an AI crawler](https://lantad.co/glossary/ai-crawler) receives, which makes them the part of [AI visibility](https://lantad.co/glossary/ai-visibility) an outside request can still settle for you. The earlier work in this field, [the GEO paper that introduced the benchmark this study samples from](https://arxiv.org/abs/2311.09735), was submitted on 16 November 2023 and its abstract states that its semantic methods can boost visibility by up to 40 percent in generative engine responses. That is a different metric from citation rate and an upper bound rather than a mean, so the two figures do not sit on one scale and neither is evidence about the other. What can be said about structure without comparing them is that it is the lever which does not require rewriting anything.

## Questions and answers

**Does content structure affect whether AI search engines cite a page?**

One study reports that it does. arXiv:2603.29979, posted on 31 March 2026, edited the structure of 200 articles without rewriting their content and reports citation rate rising from 45.0 percent to 52.8 percent across six generative engines, a gain of 7.8 percentage points. That is a single paper on 200 documents with no per-platform breakdown, so it is evidence rather than settled fact.

**Is the 17.3 percent citation gain a rise of 17.3 percentage points?**

No. The paper's Table IV records the move as 45.0 percent to 52.8 percent, which is 7.8 percentage points. The 17.3 percent figure is that gain expressed relative to the 45.0 percent baseline, and the paper's ablation table records the same distance as a drop of 7.8.

**Which structural changes did the study attribute most of the gain to?**

Macro-structure, which the paper defines as heading hierarchy, navigational elements and document flow, at 44.9 percent of the gain. Meso-structure, being paragraph organisation, lists and tables, accounts for 39.7 percent. Micro-structure, being emphasis markers, keyword placement and syntactic patterns, accounts for 15.4 percent, or 1.2 of the 7.8 points.

**Has Lantad measured any of this?**

No. Lantad runs no citation trials and holds no citation data, and this post reports the paper rather than any scan. What this scanner does measure sits earlier in the path: whether a crawler can fetch the page, and whether the text and structure it receives match what a browser renders.

---

Lantad measures whether AI crawlers can actually read a page: it fetches as a non-rendering
crawler, renders as a browser, and reports the gap. Free scan, one URL, no signup.

Method and weights: https://lantad.co/methodology | All pages as markdown: https://lantad.co/md | Crawler policy: https://lantad.co/bot
