BlogFindings

How to measure GEO when the final prompt holds 35.6 percent of the vocabulary

A paper posted to arXiv on 24 July 2026 measured where a conversation's request state actually sits across its user turns. In 670 commercial multi-turn conversations the final prompt carried a median 35.6 percent of the session's unique user-side content vocabulary, and in 50.3 percent of them a rule set detected at least one request-state dimension in the history that the final prompt never repeats.

16 min read Lantad

The finding is uncomfortable for the whole category, whether it calls itself generative engine optimization or answer engine optimization, and it is uncomfortable here first, so this post reports it and then says where Lantad sits in it. Lantad has measured nothing about multi-turn conversations. It has no conversation corpus, it has never run a session-level test, and every figure below belongs to the paper or to the dataset the paper reuses. What this site can add is the part it does know: what its own tracker sends, what that makes it under the paper's own taxonomy, and which numbers on a report survive the finding intact.

In short

  • A paper by Benjamin Tannenbaum posted to arXiv on 24 July 2026, arXiv:2607.22392v1, reports that the final prompt of a multi-turn conversation carried a median 35.6 percent of the session's unique user-side content vocabulary in a 670 conversation commercial corpus, and 36.4 percent in 7,463 public PRISM conversations.
  • The same paper reports that the final prompt held at most half of that vocabulary in 68.4 percent of the commercial conversations and 74.3 percent of the PRISM conversations, so a majority of sessions put most of their words somewhere other than the last thing typed.
  • Transparent rules detected at least one request-state dimension present in the history but absent from the final prompt in 50.3 percent of commercial conversations and 44.8 percent of PRISM conversations, and the final prompt reproduced the full observed dimension set in only 26.1 percent and 26.2 percent.
  • Anyone trying to measure GEO by replaying a single prompt is running the paper's isolated final turn condition, which is the condition it says omits observed request evidence. Lantad's own tracker is in that position: worker/src/engines.ts sends every engine a single user message with no history, read on 28 August 2026.
  • The paper is explicit about its own limit and Lantad has measured none of this: in its words, no prompt was rerun, so the results show that isolated prompts omit observed request evidence and do not show how much a model answer would change when history is removed.
What was countedCommercial corpus, 670 conversationsPRISM corpus, 7,463 conversations
Median share of the session's unique user-side content vocabulary held by the final prompt35.6 percent36.4 percent
Conversations where the final prompt holds at most half of that vocabulary68.4 percent74.3 percent
At least one request-state dimension in history but not in the final prompt50.3 percent44.8 percent
Final prompt reproduces the full observed dimension set26.1 percent26.2 percent
Final prompt adds a dimension not observed earlier17.9 percent19.3 percent
Headline results as stated in the abstract of arXiv:2607.22392v1, posted 24 July 2026 and read at arxiv.org on 28 August 2026. Both corpora are described in the paper. Lantad measured none of these values and holds no conversation data of any kind.

What the study measured, and on how many conversations

The paper replaces intent, which nobody can observe, with something that can be counted. It calls the construct conversation-conditioned request state: the constraints and decision moves a user has actually put on the record across their turns. Then it measures how that state is distributed across the turns of a session, and in particular how much of it is present in the last turn, which is the turn every prompt-replay measurement treats as the query.

Two corpora carry the result. The first is commercial and is described as a discovery-replication design: a discovery export yielding 302 English multi-turn conversations, and a non-overlapping replication source drawn from 43 further governed exports yielding 368, which pool to the 670 the abstract reports. The second is public. It is drawn from the PRISM Alignment Dataset, published on arXiv on 24 April 2024, which maps 1,500 participants from 75 countries to 8,011 live conversations with 21 large language models. The paper retains conversations with at least two user turns and reports excluding 488 that failed an English designation and 60 carrying personally identifiable information flags, which leaves the 7,463 conversations from 1,389 participants the abstract names. Those two subtractions land exactly on the stated total, which is a small thing but it is the kind of small thing worth checking before quoting a figure.

The detection method is deliberately dull, and the paper's word for it is transparent. Nine request-state dimensions are matched by case-insensitive patterns rather than by a model: price or budget, location or proximity, persona or use case, attribute requirement, time, alternatives, correction or redirect, comparison or evaluation, and explanation or evidence. The first five are constraints a user states. The last four are decision moves: naming an alternative, correcting an earlier assumption, asking for a comparison, asking for evidence. A rule fires only on explicit language, which means the counts understate rather than overstate what a conversation contains, and the paper says so.

That choice matters for how much weight the result can carry. A model-based classifier would find more state and would be impossible to audit. A pattern rule finds less and can be reproduced. When a rule set this conservative still finds a dimension in the history and not in the final prompt in half of conversations, the conservatism is working in the finding's favour rather than against it. This is a different kind of evidence from the citation studies this site usually reports, such as the survey of 45 GEO studies, because it measures the input side rather than the answer.

CorpusSourceFilter appliedConversations analysed
Commercial, discoveryA governed exportEnglish, at least two user turns302
Commercial, replication43 further governed exports, non-overlappingEnglish, at least two user turns368
Commercial, pooledThe two cohorts combinedAs above670
PRISM8,011 released conversations, 1,500 participants, 21 modelsTwo user turns minimum, less 488 non-English and 60 with PII flags7,463
How the two corpora were assembled, as described in arXiv:2607.22392v1, read at arxiv.org on 28 August 2026. The PRISM totals are cross-checked against the PRISM Alignment Dataset paper, arXiv:2404.16019, published 24 April 2024. Reported, not measured by Lantad.

How do you measure GEO if the prompt is not the query?

Take the shape of a real commercial session. Somebody opens with a broad question, the assistant answers, and the person then adds a budget. Two turns later they rule out an option they had been considering. Later still they correct an assumption the assistant made about where they are. The last thing they type might be four words asking for a recommendation, and those four words are what a prompt-tracking product would store, replay and count.

The paper's numbers say how much that last turn is carrying. A median of 35.6 percent of the session's unique user-side content vocabulary in the commercial corpus, and 36.4 percent in PRISM. In 68.4 percent of commercial conversations and 74.3 percent of PRISM conversations the final prompt holds at most half of it. The paper is careful here in a way that matters: it runs length-matched nulls and reports that low lexical coverage is largely a consequence of turn length, so the vocabulary result is interpreted as information availability rather than as semantic drift. Users are not changing the subject. Their last turn is simply short, and short turns cannot carry what earlier turns established.

The categorical result is the one that bites, because it is not a function of length. In 50.3 percent of commercial conversations and 44.8 percent of PRISM conversations the rules detect at least one request-state dimension in the history that is not in the final prompt. Among conversations that carry any dimension at all, the final prompt reproduces the full observed set in only 26.1 percent and 26.2 percent. And the endpoint is not merely a compressed summary of what came before: in 17.9 percent and 19.3 percent it adds a dimension nobody had seen earlier in the session, which the paper reads as the final turn being another state update rather than a conclusion.

For anyone building or buying a measurement, that reframes the question. A replayed prompt is not a smaller version of the session. It is one turn out of a sequence, missing about half the stated constraints in half the cases, and occasionally introducing one of its own. This sits alongside the other thing already known about this method, which is that the answer moves even when the prompt does not: a study from the University of St. Gallen found that a per-brand visibility rate from a single run carried a standard error of 0.370 and needed seven runs to fall below 0.10. Sampling noise and unit error are separate problems, and running a prompt more times does nothing about the second one. The neighbouring question of how many prompts a stable estimate needs in the first place has its own answer and its own arithmetic, and neither of them reaches this.

Sample Illustrative, not a measurement of any real site.

Where request state sits in a session, and which part a prompt-replay measurement observes. Drawn from the construct and the three comparison conditions described in arXiv:2607.22392v1, read at arxiv.org on 28 August 2026. Illustrative of the paper's design, not a measurement.

Half the conversations state something the final prompt never repeats

It is worth being precise about what a missing dimension is, because the phrase sounds abstract and the thing itself is not. If a user said in turn two that their budget is under a hundred pounds and the final turn does not mention money, then a measurement that replays the final turn is asking an engine an unconstrained question and scoring the answer as though it were the constrained one. If the user ruled out a named alternative in turn three, the replayed question still has that alternative in play. The engine is answering a different question, correctly.

This is the same class of error this site published about its own product on 22 August 2026, when Lantad's prompt tracker reported a 12.6 percent brand mention rate and the truth was near 1 percent because a name in an answer is presence and presence is not identity. Both are unit errors: something countable was counted correctly and stood for something it did not stand for. That is the failure mode a measurement product should be most afraid of, because it produces a number that is stable, reproducible and wrong, and stability reads as accuracy.

The paper's nine dimensions split usefully for this purpose. Five are constraints, and a constraint dropped from the replay makes the question broader than the user's. Four are decision moves, and those behave differently: a request for evidence or a comparison changes what kind of answer is wanted rather than what is being asked about. A session where the user has twice asked for sources is a session where the engine has been pushed towards citing, and replaying only the last turn removes that pressure. Since citing is the entire surface this category measures, that is not a marginal detail. It connects to what is already known about how citation behaves under different conditions, including that more citations did not mean more of your page reached the answer text that topical relevance and list position drove the first citation while formatting did not, and that a brand's own domain drew 2.9 percent of the citations in one measurement.

None of this makes single-prompt measurement worthless. It makes it a measurement of a specific and narrow thing: what an engine says to a cold, unconstrained question of that shape. That is a real question with real value, particularly for a brand checking whether it appears at all in a category answer, which is roughly what our own prompt tracking surface and the public answer view report. The error is only in reading it as a measurement of what buyers experience, because buyers arrive at their final turn having said four other things first. The downstream half of that experience is measured separately again, and citations reached 6.8 percent of ChatGPT prompts with the visit landing on the homepage, which is a third unit in the same chain.

  • Dimension in history, absent from final prompt, commercial 50.3% 670 conversations
  • Dimension in history, absent from final prompt, PRISM 44.8% 7,463 conversations
  • Final prompt adds an unseen dimension, commercial 17.9% The endpoint is another state update
  • Final prompt adds an unseen dimension, PRISM 19.3% The endpoint is another state update
  • Final prompt reproduces the full dimension set, commercial 26.1% Among dimension-bearing conversations
  • Final prompt reproduces the full dimension set, PRISM 26.2% Among dimension-bearing conversations
Share of conversations in which the rule set found at least one request-state dimension in the history but not in the final prompt, and share in which the final prompt reproduced the full observed dimension set. Figures from the abstract of arXiv:2607.22392v1, read at arxiv.org on 28 August 2026. Not a Lantad measurement.

What our own tracker sends, and what that makes it

The honest thing to do with a finding like this is to check where your own product stands in it before recommending anything to anybody, so here is that check, read out of this repository on 28 August 2026 rather than remembered.

Lantad's prompt tracker sends each engine a single user message and no history. In worker/src/engines.ts the request body for every engine on the ladder is built as a messages array holding one object with a user role and the prompt as its content. There is no prior turn, no assistant reply, no accumulated constraint. Under the paper's own taxonomy that is precisely the isolated final turn, which is one of the three conditions it says a causal study would need to compare, the other two being full interleaved history and user-turn-only history. Two further constants in core/src/config.ts describe the rest of a run: PROMPT_RUN_GENERATION_CALLS is 2, because an unparseable candidate reply gets exactly one stricter retry before templates take over, and PROMPT_RUN_CONTROL_CALLS is 1, the branded control asked once per run.

So the accurate description of what this product measures on that surface is: what named engines say to a single, cold, unconditioned prompt, repeated on a schedule. That is a defensible thing to sell as long as it is described that way, and it is the same discipline this site applies to a grade it cannot support, which the methodology page sets out and the withheld-grade rule sits next to. It is not a defensible thing to sell as a measurement of how an engine treats your brand in the conversations your buyers actually have, and nothing in this post should be read as saying the two are close.

There is a limit on the other side too, and it protects nobody to leave it out. The paper measured conversations between people and assistants. It did not measure the engines this or any other tracker queries, it did not measure whether a session-level query produces a different set of cited domains, and it did not measure any tracking product. Whether the share of voice numbers a category of tools reports would move under a session-level protocol is an open question that this paper does not answer and that Lantad has not tested.

Sample Illustrative, not a measurement of any real site.

What the session put on the record

  • Turn 1: an open question about a category
  • Turn 2: a budget stated (price or budget)
  • Turn 3: an option ruled out (alternatives)
  • Turn 4: a location corrected (correction or redirect)
  • Turn 5: a short request for a recommendation

What the tracker sends

  • messages: [{ role: "user", content: prompt }]
  • No assistant turns
  • No prior user turns
  • No stated budget
  • No ruled-out alternative
The isolated final turn against the session it came from. The right panel is the message shape built in worker/src/engines.ts, read on 28 August 2026. The conversation on the left is constructed for this post to show the paper's dimensions and is not a real user session.

What the study does not establish, and what to change in a measurement you run

The paper puts its own boundary in one sentence, and it is the sentence to quote if only one is quoted: no prompt was rerun, the results show that isolated prompts omit observed request evidence, and they do not show how much a model answer would change when history is removed. That is a gap between omission and consequence, and it is a wide one. It is entirely possible that engines recover most of the missing constraint from context they infer, or that the answers converge anyway. Nobody has shown that either.

The paper lists further limits and they are the ordinary ones for observational work, stated plainly. Observable cues cannot reach latent intent, only what a user wrote down. Lexical coverage depends on turn length, which is why the vocabulary figure is read as information availability. Conversation depth is endogenous rather than assigned, so longer conversations may differ in kind from short ones. The commercial data is not a random sample. The assistant's own language may shape the user's, and that influence is not identified.

What the paper recommends is procedural rather than dramatic: measure at the session level, preserve turn boundaries, version the history policy, report which state dimensions were available at generation time, and record which constraints were established earlier rather than in the final turn. For a causal claim it names the three conditions to compare, being full interleaved history, user-turn-only history, and the isolated final turn. Nothing there requires a new model or a new dataset. It requires writing down which of the three you ran, which almost nobody does today.

Three practical consequences follow for anyone reading a visibility number this week. First, ask what unit produced it. A percentage built from isolated prompts is a real measurement of a narrow thing and should be labelled as one, in the same way Search Console's generative AI report counts impressions and nothing else and is useful precisely because it says so. Second, do not compare across tools that disagree about the unit, which is the same trap as treating two engines' surfaces as interchangeable, covered here in the note that a Gemini citation tool is not an AI Overview tool. Third, keep the measurement you have and describe it accurately, because the alternative on offer is not a better number, it is an unbuilt protocol.

One thing this finding leaves entirely alone is worth saying last, because it is the part most people can act on. Everything above concerns what an engine is asked. None of it touches whether the engine can read your pages when it goes looking, which is a separate layer with its own failure modes, its own crawler vocabulary and its own per-platform guidance such as how to get cited by ChatGPT. A page that cannot be fetched is invisible under every measurement protocol anyone has proposed, single turn or session level, so nothing here is a reason to postpone that work. The paper itself is worth reading in full at arxiv.org/abs/2607.22392.

  • Measure at the session level The tracker sends one user message per engine call, so every run is the isolated final turn condition.
  • Preserve turn boundaries There are no turns to preserve: nothing prior is sent with the prompt.
  • Version the history policy The policy is fixed and readable in worker/src/engines.ts rather than configurable per run, so it cannot drift unrecorded.
  • Report state available at generation time Not reported, because no state beyond the prompt itself exists in a run.
  • Record which constraints were established earlier Not applicable to a single-turn protocol, and that is the point rather than a defence.
  • Say which condition the number came from This post and the methodology page say it. The paper's complaint is that most measurements do not.
The measurement practices arXiv:2607.22392v1 recommends, read at arxiv.org on 28 August 2026, against what Lantad's prompt tracker does today, read out of worker/src/engines.ts and core/src/config.ts on the same date.

Written by

Lantad

Published .

Every product that sells AI visibility measurement, this one included, works by sending a prompt to an engine and reading what comes back. The prompt is treated as the unit: something you can write down, count, classify, store and replay next week to see whether your number moved. A paper posted to arXiv on 24 July 2026 by Benjamin Tannenbaum, titled The Prompt Is Not the Query, takes that unit apart. It does not ask whether the answers are right. It asks whether the last thing a person typed is a reasonable stand-in for what they were actually asking, and it answers with counts rather than argument.

Common questions

How do you measure GEO if a single prompt is not the query?

The paper posted to arXiv on 24 July 2026 recommends session-level measurement: preserve turn boundaries, version the history policy, report which request-state dimensions were available at generation time, and record which constraints were established in earlier turns rather than the final one. For a causal claim it names three conditions to compare, being full interleaved history, user-turn-only history, and the isolated final turn. It does not publish a tool or a protocol implementation, so this is a description of what to record rather than software anyone can run today.

Does this mean single-prompt AI visibility numbers are wrong?

No, it means they measure a narrower thing than they are usually read as measuring. A single-prompt number is an accurate measurement of what an engine says to a cold, unconstrained question of that shape. The paper's finding is that a real session usually has more on the record than that: a median 35.6 percent of the session's unique user-side content vocabulary sat in the final prompt in its 670 conversation commercial corpus, and at least one request-state dimension was present in history but absent from the final prompt in 50.3 percent of them.

Has Lantad measured any of this?

No. Lantad holds no conversation corpus, has run no session-level test, and every figure in this post belongs to arXiv:2607.22392v1 or to the PRISM Alignment Dataset it reuses. The only claims here read out of this repository are about Lantad's own code: worker/src/engines.ts builds each engine request as a single user message with no history, and core/src/config.ts sets PROMPT_RUN_GENERATION_CALLS to 2 and PROMPT_RUN_CONTROL_CALLS to 1, both read on 28 August 2026.

What is a request-state dimension?

It is one of nine categories the paper detects with case-insensitive patterns rather than with a model, so the detection can be audited. Five are constraints a user states: price or budget, location or proximity, persona or use case, attribute requirement, and time. Four are decision moves: alternatives, correction or redirect, comparison or evaluation, and explanation or evidence. A rule fires only on explicit language, so the counts are conservative by construction.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.