BlogFindings
Four factors decided the first citation, and formatting was not one
A controlled study posted to arXiv on 25 May 2026 ran 252,000 trials across six language models, showing each of them two competing sources that differed in exactly one of 18 content factors. Four factors moved the first citation in every model, one of them was list position, and the two formatting factors are grouped by the paper among the seven with no consistent effect.
What Gets Cited: Competitive GEO in AI Answer Engines was posted to arXiv on 25 May 2026 by Rahul Vishwakarma, Shushant Kumar and Ratnesh Jamidar, all three at Sprinklr, under CC BY 4.0 and accepted to the 49th ACM SIGIR conference. It runs 252,000 trials across six language models to answer one narrow question: when two retrieved sources compete for the same citation slot, what makes one of them the first citation in the answer? The reported result is a hierarchy. Four factors moved the outcome in every model tested, one of those four is list position rather than anything written on the page, and the two formatting factors sit in the group the paper says had no consistent effect. Lantad did not run this experiment, operates no engine tracking and holds no citation corpus, so every figure below is reported from the paper. What this site can add is the condition the apparatus assumed and a site owner has to earn: every trial handed the model the full text of both pages before the comparison began.
In short
- A study posted to arXiv on 25 May 2026 as arXiv:2605.25517, by three authors at Sprinklr and accepted to SIGIR 2026, ran 252,000 trials across six language models to measure which of two competing sources an answer engine cites first.
- The paper names four gatekeeper factors that reached significance in all six models: topic match, an explicit price, a recent timestamp against an old one, and list position.
- The paper describes those four as carrying large effects, while its own Table 2 reports odds ratios for them ranging from 6.26 to more than 10,000, so the unanimity is about direction across models rather than about every cell clearing one threshold.
- For structured against dense formatting, the same table reports odds ratios of 1.68, 1.03, 0.79, 0.90, 0.78 and 1.25 across the six models, and the paper's own scale calls anything below 1.5 negligible.
- Lantad ran none of these trials and holds no citation data. Every trial in the study handed the model the full text of both sources through a simulated tool response, and the paper states that no search engine was ever called, so the experiment begins after the fetch that Lantad measures.
| Factor, stronger variant first | Lowest reported odds ratio | Highest reported | What differed between the two sources |
|---|---|---|---|
| On-topic against off-topic | 221 | > 10,000 | One source discussed unrelated products |
| Price against no price | 6.26 | > 10,000 | One source omitted the product price |
| Recent against old timestamp | 14.4 | > 10,000 | Content dated 2026 against content dated 2019 |
| Position 1 against position 2 | 1,795 | > 10,000 | Nothing on the page. The order the two sources were listed in |
What a two source citation test can settle
The apparatus is the reason the numbers are worth quoting, and it is also the reason they have a narrow scope. The paper builds what it calls a two document retrieval augmented generation testbed. For each trial the model receives a question, then a simulated web search tool call, then a tool response listing exactly two candidate sources, each with a title, a URL and the full text. The model answers with citations, and the recorded outcome is which of the two URLs appears first in the answer. The paper states plainly that it never calls a search engine, and that the same tool call and query string are reused on every repeated run of a pair.
The corpus was built to remove two confounds that make observational studies of AI answers hard to interpret. First, familiarity: the authors used GPT-4o to curate 100 product review blog posts across 50 consumer product categories, then to replace every brand name, product model and publisher name with a fictional alias while preserving prices, specifications and timestamps. Second, order: for 17 of the 18 factors the two sources were shown in both orders, so a preference for whatever appears first could be separated from a preference for the content. The eighteenth factor is list position itself, where order is the treatment and so is not counterbalanced.
The scale follows from that design. For each of the 18 factor hypotheses the authors selected 20 blogs and generated four scenarios each, giving 80 scenarios per factor and 1,440 in total, where the two variants differ in exactly one factor and match on facts, prices, specifications and length to within five percent. Three query paraphrases per scenario produced 4,320 scenario and query instances, each run five times, across six models: Gemini-2.5-Flash, Claude-3.5-Sonnet, Kimi-K2-Thinking, GPT-5-Nano, GPT-5-Mini and GPT-5.2. That is 2,400 trials for each of 17 factors, 1,200 for the position factor, 42,000 per model, and 252,000 in total. All three authors independently reviewed a stratified sample of 300 scenarios, 21 percent of the 1,440, and report finding no brand leakage and no unintended factor differences.
This design answers a question that observational counts cannot. Watching live answer engines mixes content quality with retrieval rank, interface behaviour and brand familiarity, which is why measuring visibility from live prompts needs so many repeated queries before the number settles. It also answers a narrower question than the phrase AI visibility usually covers, because the two candidates are already in the model's context when the trial starts.
The authors are explicit about where their question sits relative to the work that named the field. They follow the 2023 GEO paper by Aggarwal and colleagues, accepted to KDD 2024, in calling this generative engine optimization, and state their departure from it: that paper quantified single source visibility within a fixed retrieved set, while this one estimates citation preference, meaning which of two similar candidates wins when they compete directly for the same slot.
Flow: 100 review blogs, 50 categories to Brands replaced with aliases; Brands replaced with aliases (one factor differs) to 1,440 one factor pairs; 1,440 one factor pairs (3 paraphrases) to 4,320 scenario queries; 4,320 scenario queries (both orders) to 6 models, 5 runs each; 6 models, 5 runs each to 252,000 trials; 252,000 trials to First URL in the answer.
The four factors every model agreed on
Of the 18 factors, the paper reports that 11, or 61 percent, reached significance in four or more of the six models. Four of those it separates out and calls gatekeepers, on the grounds that all six models agreed on them: topic mismatch, price not mentioned, a recent timestamp against an old one, and lower list position. Its summary sentence is that failing on any one of them can eliminate citation odds regardless of other content strengths.
Two of the four are worth pausing on because they are not the levers the GEO advice market tends to sell. An explicit price is a piece of information, not a formatting choice, and the corpus was built from consumer product reviews where a price is the natural thing a reader wants. A timestamp is also information rather than structure. Neither requires a schema block, a heading rewrite or a content management system change, and both are the sort of thing that goes missing when a page is written to be evergreen. This lines up with the shape of the five GEO tactics Google's own guide tells site owners they do not need, where the advice repeatedly comes back to the page saying something concrete rather than being shaped a particular way.
There is a discrepancy inside the paper worth naming, because reading the abstract alone would hide it. The results section describes the four gatekeepers as unanimous across all six models with large effects, giving an odds ratio above 100 as the marker. Its own Table 2 does not support that reading cell by cell. For price against no price the table reports 36.1 for Kimi-K2-Thinking, 7.82 for GPT-5-Nano, 6.26 for GPT-5-Mini and 30.4 for GPT-5.2, all below 100 and, on the paper's own effect size scale, in the moderate to strong bands rather than the very strong one. For recent against old timestamp it reports 68.7 and 14.4 for two of the models. The unanimity the paper is describing is agreement on direction and significance across all six models, which is a real and useful property, rather than a claim that every cell clears one threshold.
The paper is also explicit about why so many cells read as greater than ten thousand. Where preference is nearly deterministic the fitted odds ratio becomes unstable under quasi-separation, and the authors say they report those as greater than 10k and treat them as a decisive win rather than a finely resolved ratio. That is the correct way to present the number, and it means the largest values in the table should be read as one bit of information rather than as a magnitude. Anyone reusing the table for a structured data argument or a content brief should carry that caveat with the figure.
| Factor | Gemini 2.5 Flash | Claude 3.5 Sonnet | Kimi K2 Thinking | GPT 5 Nano | GPT 5 Mini | GPT 5.2 |
|---|---|---|---|---|---|---|
| On-topic vs off-topic | > 10k | > 10k | > 10k | 221 | > 10k | > 10k |
| Price vs no price | > 10k | > 10k | 36.1 | 7.82 | 6.26 | 30.4 |
| Recent vs old timestamp | > 10k | > 10k | 68.7 | 14.4 | 1,494 | > 10k |
| Position 1 vs position 2 | > 10k | > 10k | > 10k | 2,002 | 1,795 | > 10k |
| Specs vs no specs | 8.63 | > 10k | 238 | 11.5 | 15.4 | 243 |
| Consistent vs contradictory | 2.81 | 2.81 | 2.72 | 2.19 | 4.09 | 1.74 |
One of the four gatekeepers is not on your page at all
Lower list position is the factor the study treats differently from the other 17, and it is the one with the most awkward implication for anybody selling page level optimisation. In this test it is the only factor where the two variants are identical in content: the same text appears as source one or as source two, and nothing else changes. The odds ratios reported for position one against position two are greater than 10k for Gemini-2.5-Flash, Claude-3.5-Sonnet, Kimi-K2-Thinking and GPT-5.2, 2,002 for GPT-5-Nano and 1,795 for GPT-5-Mini.
Put next to the content factors, that is the study saying something uncomfortable about the whole exercise. The largest and most consistent effect it measured came from where a source sat in the retrieved list, which is decided by retrieval before the model sees anything. The paper does not treat this as a flaw. It cites the established position bias literature as the reason to expect it, and its practitioner workflow is explicit that if a brand is absent from citations entirely the bottleneck is retrieval rather than content.
That separation matters when reading any citation statistic. Citation counts from live engines are a joint product of retrieval and selection, which is why more citations did not mean more of the page reached the answer in the measurement framework we read earlier this month, and why the destinations of AI citations look so different from a site owner's expectation once the classes of cited source are broken out. A study that isolates content factors has to hold retrieval fixed, and this one does so by injecting the candidates directly. The cost of that choice is that it cannot say anything about how a page gets into the list in the first place.
The honest reading for a site owner is therefore a sequence rather than a checklist. The GEO work this study measures applies to pages that are already being retrieved as candidates. For a page that is never retrieved, the factors in Table 2 are not the binding constraint, and no amount of adding prices and dates will make them so.
-
Gemini 2.5 Flash> 10k Reported under quasi-separation, meaning a decisive preference for the first listed source -
Claude 3.5 Sonnet> 10k Reported under quasi-separation -
Kimi K2 Thinking> 10k Reported under quasi-separation -
GPT 5 Nano2,002 A finite odds ratio, still the largest or near largest effect for this model -
GPT 5 Mini1,795 A finite odds ratio, still far above any content factor for this model -
GPT 5.2> 10k Reported under quasi-separation
Formatting was the one lever that did not move the citation
Seven of the 18 factors, 39 percent, are grouped by the paper as having weak or no effect. Two of those seven are the formatting factors. The results section states that formatting choices, naming content structure and scattered information, had no impact, and offers the interpretation that the models parse content regardless of visual organisation.
The table is the more useful record here because it shows the size of what was measured rather than a verdict. For structured against dense formatting, where one variant is organised into sections and the other is a dense paragraph block, the reported odds ratios are 1.68 for Gemini-2.5-Flash, 1.03 for Claude-3.5-Sonnet, 0.79 for Kimi-K2-Thinking, 0.90 for GPT-5-Nano, 0.78 for GPT-5-Mini and 1.25 for GPT-5.2. Three of the six sit below 1.0, which is the direction opposite to the one the hypothesis predicted, and the paper's own effect size scale calls anything below 1.5 negligible. For organised against scattered information the reported values run from 1.13 to 3.87, a wider spread, and the paper still places that factor in the group without a consistent effect because its rule for consistency is agreement across models rather than the size of any single cell.
This is the finding that is least convenient for the category this site sits in, so it is worth stating without softening. Lantad scores page structure, and this study offers no evidence that reorganising text a model can already read wins a citation against a competing source. The two are not measuring the same thing: our scoring method treats structure as part of whether the text a human reads is present and extractable in the served response, while this study varied the arrangement of text that had already been placed in the model's context in full. But the distinction does not rescue the broader claim that headings and section breaks are a citation lever, and nothing in this table supports that claim.
The pattern is familiar from other controlled work on AI search. When five attack modes were tested against retrieval augmented answers, hidden text planted in the machine readable page was the weakest of them at an average success rate under one percent, while fabricated agreement between sources was the strongest. Both results point the same way: these systems respond to what the text asserts and how well it matches the question, not to how the text is dressed. That is also why prose parity is a threshold rather than a score to maximise. Text that is absent cannot assert anything, and text that is present is apparently read whatever shape it is in.
| Factor | Gemini 2.5 Flash | Claude 3.5 Sonnet | Kimi K2 Thinking | GPT 5 Nano | GPT 5 Mini | GPT 5.2 |
|---|---|---|---|---|---|---|
| Structured vs dense | 1.68 | 1.03 | 0.79 | 0.90 | 0.78 | 1.25 |
| Organised vs scattered | 2.21 | 3.87 | 2.49 | 1.19 | 1.13 | 1.57 |
The six models did not agree about much beyond those four
Underneath the headline hierarchy the models behave differently enough that a single ranked checklist hides most of the variation. The paper reports Kimi-K2-Thinking as the most content sensitive, with 83 percent of the factors reaching significance, followed by the GPT family, then Claude-3.5-Sonnet at 50 percent and Gemini-2.5-Flash at 33 percent. It gives no per model figure for the three GPT models individually, so neither do we.
It also reports a difference in the shape of the responses rather than just their number. Gemini and Claude are described as showing categorical patterns, with 67 to 78 percent of their significant factors producing odds ratios above 10,000, which the authors read as a sign of near binary decision boundaries. The GPT models, by contrast, are described as behaving consistently across the three sizes tested, which the paper takes as evidence that citation behaviour is shaped by architecture rather than model scale.
One further number is worth carrying because it constrains how the outcome was defined. Across successful runs, answers contained exactly one distinct URL about 86.4 percent of the time, two or more about 10.5 percent, and no URL at all about 3.1 percent. Roughly 3.4 percent of trials were dropped from the regression because they carried no URL or a first URL matching neither injected variant. So the modal answer in this testbed cited one source out of two, which is the competitive condition the study set out to create, and it is a sharper condition than a live answer that cites several pages.
The practical consequence is that a result stated as a property of AI answer engines in general is really a property of six specific models on one date. That is not a criticism of the study, which names its models and its scale, but it is a reason to be careful about how the finding travels. Guidance written for one assistant does not transfer intact to another, which is why the advice for getting cited in ChatGPT is a separate page here from the advice for the others, and why our own research page reports what the sample supports rather than a single universal rule.
What has to be true before any of this applies to your page
The single most important sentence in the paper for a site owner is in its methodology rather than its results: the tool response always lists exactly two variant sources, each with a title, a URL and the full text, and no search engine is ever called. Every factor in Table 2 was measured under the assumption that the complete text of both candidate pages was already in front of the model. That assumption is the thing a real page has to earn, and it is not free.
The study is honest about the boundary. Its limitations section notes that production retrieval augmented systems often retrieve five to ten or more pages, so real citation pools are larger than a pair, and that the estimates are pairwise preferences over a controlled slate rather than full multi document competition. It notes that brands and publishers were anonymised, so any residual preference a production system holds for a trusted domain is outside what was measured. It notes that GPT-4o generated the seed corpus, the anonymisation and the paired rewrites. And its own practitioner workflow routes the case where a brand appears in no citation at all away from content work entirely, to retrieval.
That upstream step is what an external readability scan settles, and it fails for ordinary reasons that have nothing to do with content quality. A page whose text arrives only after JavaScript runs is a different document to a client that does not execute it, which is why the July Common Crawl archive contains no JavaScript execution and why the served HTML and the rendered page have to be compared rather than assumed equal. A robots.txt rule resolved against one named AI crawler token can allow a fetch that another token is refused, which is one half of the two layers that decide whether AI can read your site. And these are not rare: across the six real site captures stored on 15 July 2026, two sites on the same platform produced opposite outcomes on the same day.
None of that is a claim about citation. Lantad has never tested whether readability correlates with citation rate on any cohort, holds no citation data, and runs no prompts against any engine, so the relationship between the two remains unmeasured here. What can be said is a scope statement rather than a finding. This study measured what changes a preference between two documents already in context. A scan measures whether your document reaches that context at all: whether the named token is allowed, what status code and bytes a client without a browser receives, how much of the visible text is present before any script runs, and what the structured data looks like on arrival. If you want to see the second of those for yourself, what GPTBot sees renders the response a crawler receives, and the prose parity comparison behind it is documented. The paper measured the second half of the problem carefully. The first half is still yours.
Supplied by the testbed
- Exactly two candidate sources, injected directly into the model context
- A title, a URL and the full text of each page
- No robots.txt lookup, because no search engine was ever called
- No rendering step, because no page was ever fetched
- Brands and publishers replaced with fictional aliases
What a real page has to clear first
- A robots.txt rule resolved for the specific named crawler token
- A status code and a response body returned to a client with no browser
- The visible text present before any JavaScript executes
- Structured data parseable in the response as served
- Being retrieved as a candidate at all, which this study held fixed
Lantad
Published .
Most advice about generative engine optimization arrives as a list of page level edits: add headings, add schema, rewrite in a question and answer shape, refresh the date. The list is rarely accompanied by a measurement of which edit moved anything, and the survey of 45 GEO studies we read in July reported that none of the studies it reviewed shows a stable, cross-platform causal effect on organic discoverability. A controlled experiment that ranks those edits against each other, rather than testing one of them in isolation, is worth reading at the source.
Common questions
What makes an AI answer engine cite one page rather than another?
In this study, four factors moved the outcome in all six models tested: whether the page is on topic for the question, whether it states an explicit price, whether it carries a recent timestamp rather than an old one, and where it sits in the retrieved list. The study ran 252,000 trials and was posted to arXiv on 25 May 2026 as arXiv:2605.25517. Its testbed showed each model exactly two competing sources with the full text of both already supplied, so the result describes selection between candidates rather than how a page becomes a candidate.
Does adding headings and structure help a page get cited by AI?
Not in this experiment. The paper groups content structure and scattered information among the seven of 18 factors with no consistent effect, and its Table 2 reports odds ratios for structured against dense formatting of 1.68, 1.03, 0.79, 0.90, 0.78 and 1.25 across the six models, three of which point in the opposite direction to the hypothesis. The paper's own effect size scale calls anything below 1.5 negligible. This is a finding about rearranging text a model already holds, not about whether the text is present in the response a crawler receives.
Why does list position matter more than anything written on the page?
Position was the largest and most consistent effect the study measured, with odds ratios for position one against position two of more than 10,000 in four models and 2,002 and 1,795 in the other two. It is the only factor where the two variants were identical in content. Position is decided by retrieval before the model reads anything, and the paper's own practitioner workflow routes a brand that appears in no citation at all to retrieval work rather than content work.
Did Lantad measure any of this?
No. Lantad ran none of these trials, operates no engine tracking, holds no citation corpus and has never tested whether crawler readability correlates with citation rate. Every figure in this post is read from arXiv:2605.25517v1 on 12 August 2026. What Lantad measures is the step before the one this study tests: whether a named AI crawler is allowed by robots.txt, what a client without a browser receives, and how much of the visible text is present before JavaScript runs.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.