Blog / A survey of 45 GEO studies puts crawling second and says few observe it

A survey of 45 GEO studies puts crawling second and says few observe it

A critical survey posted to arXiv on 15 July 2026 reviewed 45 generative engine optimization studies and reported that none of them shows a stable, cross-platform causal effect on organic discoverability. Its own pipeline diagram puts crawling and indexing at stage two of seven, and its caption says far fewer studies look there.

In short

  • Optimizing Visibility in Generative Engines, a critical survey by Olivier Martinez posted to arXiv as 2607.14035 on 15 July 2026, reviews 45 studies selected under a publication window running from 16 November 2023 to 14 July 2026.
  • The survey's conclusion states that within its reviewed corpus no technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream clicks and conversions.
  • The paper's Figure 1 places crawling and indexing at the second of seven pipeline stages, and the caption under it says far fewer studies observe crawling, organic retrieval, or user behavior than observe the stages between context allocation and citation.
  • SAGEO Arena, reported in the survey at 171,003 documents and 2,700 queries, found that body-only optimization reduced average top-20 presence by approximately 9 percent, top-10 presence after reranking by 16 percent, and final citation by 6 percent.
  • Lantad measures crawler access and what a page returns without JavaScript, which is the survey's second stage and nothing further; every figure in this post is transcribed from the paper rather than measured by Lantad.

The commercial vocabulary around AI search has run well ahead of the published evidence, and until now there has been no single place to see how far. There is one now. A critical survey of the generative engine optimization literature went up on arXiv on 15 July 2026, filed under cs.IR by Olivier Martinez as 2607.14035v1 under a CC BY-SA 4.0 licence. It reads 45 studies, grades what each one actually establishes, and states plainly where the field's claims outrun its data.

Two things in it are worth a site owner's attention. The first is the headline judgement, which is unusually blunt for a literature review. The second is quieter and more useful: the paper draws the visibility pipeline as seven stages, puts crawling and indexing at stage two, and then notes that hardly anybody studies that stage. Lantad measures stage two. It does not measure the other six, and this post does not pretend otherwise. Everything below is reported from the survey as read on 31 July 2026, with the arithmetic and the interpretation marked where they are ours.

  • Activation Rarely observed Whether the engine decides to search at all.
  • Crawling and indexing Rarely observed Stage two of seven. The stage an external scan can reach.
  • Retrieval Rarely observed Named in the caption alongside crawling as under-observed.
  • Reranking and context Heavily studied Inside the band the caption says most work optimizes.
  • Generation and citation Heavily studied The end of that band.
  • Absorption and fidelity Studied separately Whether a claim is used and whether it is supported.
  • Attention, click, conversion Rarely observed Named in the caption as user behavior.
The seven stages of the causal visibility pipeline as named in Figure 1 of arXiv:2607.14035, posted 15 July 2026. The coverage column restates that figure's caption, which says most studies optimize the stages between context allocation and citation and far fewer observe crawling, organic retrieval, or user behavior.

What the survey reviewed, and what it declines to claim

The paper describes itself as a critical scoping review rather than a systematic review in the clinical sense, and it gives its reasons: the field has not converged on a stable vocabulary, many studies exist only as preprints, and the systems under study change over the publication cycle. That is a limitation stated by the authors, not a criticism from us, and it sets the weight every figure below can carry.

The corpus is 45 studies. The primary window opens on 16 November 2023, which the paper identifies as the date of the first arXiv version of the foundational GEO paper, and closes on 14 July 2026. One earlier preprint is admitted because its EMNLP proceedings publication in December 2023 fell after that opening date. The search covered arXiv, the ACM Digital Library, the ACL Anthology, the NeurIPS proceedings, PMLR and OpenReview, followed by backward and forward citation searching from the core studies, and a study qualified if it met at least one of four stated criteria covering interventions, commercial-engine measurement, attacks and defenses, or the incentives created by the distribution of visibility.

The caution about publication status is the part most likely to be dropped when this paper gets summarised elsewhere, so it is worth stating here. The corpus mixes peer-reviewed articles, accepted but forthcoming papers, workshop papers and preprints, and the survey says that distinction is substantive rather than cosmetic. It names two studies as accepted to SIGIR 2026 and describes them as forthcoming because, as of 14 July 2026, that conference had not yet begun. It names four measurement studies that had not been peer reviewed, says it uses their data where they complement published findings, and says it does not assign them the same evidentiary weight. It also records that the original search did not retain database-specific hit counts or a complete exclusion ledger, and marks that as unavailable rather than reconstructing it.

That is a higher standard of self-disclosure than most vendor research in this category, ours included, and it is the reason this post treats the paper as citable. It is the same standard we hold our own published research to and the same one behind how we describe our method. Readers who want the adjacent question of how many prompts an answer engine optimization measurement needs before it means anything will find a different paper covered in our post on measurement sample size.

  • Studies reviewed 45 Plus relevant RAG and evaluation work, per the abstract.
  • Databases and archives searched 6 arXiv, ACM DL, ACL Anthology, NeurIPS, PMLR, OpenReview.
  • Inclusion criteria, any one sufficient 4 Intervention, commercial measurement, attack or defense, incentives.
  • Window, days 16 Nov 2023 to 14 Jul 2026 Opens at the first arXiv version of the foundational GEO paper.
How the corpus was assembled, transcribed from Section 2 of arXiv:2607.14035, posted 15 July 2026. Publisher figures, not a Lantad measurement.

Why the crawling stage sits second and gets studied least

Section 3.1 formalises a generative engine as a chain rather than a ranking function. The engine may first decide whether to activate search at all. If it does, it retrieves a set of documents, ranks or reranks them, and a generator then produces an answer and a set of citations from a context window. The paper's point is that content creators do not generally observe the retrieved set, the reranking score, or the generator's internal states, which makes this a black-box optimization problem under incomplete information.

Figure 1 draws that chain as seven stages: activation, crawling and indexing, retrieval, reranking and context, generation and citation, absorption and fidelity, and attention, click and conversion. The caption is the sentence worth pinning to a wall. Most GEO studies optimize the stages between context allocation and citation, it says, and far fewer observe crawling, organic retrieval, or user behavior.

Read that against the ordering and the consequence is arithmetic rather than argument. Stage two is a precondition for stages three through seven. A page that an AI crawler cannot fetch, or can fetch but cannot read because the text arrives only after a JavaScript bundle executes, does not enter the retrieval set, so no amount of work on the stages the literature does study can apply to it. The survey does not say this in so many words about any individual site. What it does say, in Section 13.1, is that the research priority is a benchmark in which a modified page must actually be crawled, indexed, retrieved, reranked and then cited, and that content modifications, internal linking, structured data, domain reputation and crawler accessibility must be separated as factors. Crawler accessibility appears there as a variable nobody has isolated.

Section 3.2 makes the same point from the other end by refusing to collapse visibility into one number. It proposes a vector of seven components: retrieval probability, context exposure, citation probability, observable prominence, absorption, fidelity, and behavioral or economic outcomes. The paper then says a high conditional probability of citation does not compensate for a low probability of retrieval, and that a scalar score is defensible only when its weights correspond to an explicit objective. That is the clearest available statement of why a single AI visibility number, ours or anyone's, has to disclose what it is made of. Anyone wanting to see the fetch side of stage two on a specific URL can run our crawler-view fetch, and the reason access is decided in more than one place is covered in the post on the two layers.

Stages most GEO studies optimize

  • Reranking and context allocation
  • Generation and citation
  • Effects measured with the document already in context
  • Table 5 grades these claims High confidence

Stages far fewer studies observe

  • Crawling and indexing
  • Organic retrieval
  • User behavior, clicks and conversions
  • Table 5 grades these Low and Very low
The division stated in the caption of Figure 1, arXiv:2607.14035. Left and right are the paper's own grouping of where research attention sits, not a Lantad assessment of any tool or site.

What the evidence supports, graded by the survey itself

Table 5 in Section 10 is the most quotable object in the paper because it grades ten claims by confidence and attaches the caveat to each. Four are marked High. A document already placed in the context can causally alter its rank, citation or use, on the basis of controlled replications across multiple models, with the explicit note that this does not address organic retrieval. Query to document relevance and context position are major determinants. Commercial engines differ from one another and vary over time. Retrieved documents constitute a genuine attack surface.

Three are Moderate: that extractable evidence and suitable structure often facilitate use, dependent on intent, engine and factuality; that systematically optimized or learned methods often outperform fixed heuristics in controlled benchmarks; and that competitive adoption can erode individual gains in tested multi-actor settings.

Then the gradient falls away. The claim that a white-hat GEO intervention durably improves organic discoverability across multiple engines is graded Low, with the caveat that there are very few end-to-end tests. The claim that citation scores predict clicks, conversions or revenue is graded Very low, resting on one suggestive quasi-experiment and a few industry claims, with causality not established. And one row is marked rejected as a general claim: the familiar figure that GEO increases visibility by 40 percent, which the survey says is a relative maximum on one metric under a specific configuration. The introduction adds that the number describes a relative gain in a simulator in which five documents have already been placed in context.

The conclusion in Section 15 states the synthesis directly. Already retrieved content can causally influence an answer, while the review identified no technique with a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream clicks and conversions. The paper is careful that this bounds a corpus rather than the world: Section 2.4 says any claim of absence is bounded to the reviewed corpus, and the conclusion adds that this limitation does not invalidate GEO but defines the work still required. Note what survives that grading, because it is not nothing. Structure and extractable evidence keep a Moderate grade, which is roughly where structured data belongs in anyone's priorities, and we have published our own inconvenient result on how differently two readings of the same markup can score in a post on measuring schema twice. Platform-specific guidance such as what it takes to be cited in ChatGPT should be read with the same grading in mind.

  • High: already-placed content can alter its rank, citation or use Controlled replications across multiple models; does not address organic retrieval.
  • High: relevance and context position are major determinants Counterfactual experiments, benchmarks and a factorial study; effects may vary with length and task.
  • High: commercial engines differ and vary over time Published audits and large preprints over heterogeneous products and periods.
  • Moderate: extractable evidence and structure often facilitate use Consistent findings, but dependent on intent, engine and factuality.
  • Low: a white-hat intervention durably improves organic discoverability Very few end-to-end tests; SAGEO even finds adverse upstream effects.
  • Very low: citation scores predict clicks, conversions or revenue One suggestive quasi-experiment and a few industry claims; causality not established.
  • Rejected as a general claim: GEO increases visibility by 40 percent A relative maximum on one metric under a specific configuration.
Table 5 of arXiv:2607.14035, Confidence in the field's principal claims, transcribed on 31 July 2026. The grades and caveats are the paper's, not Lantad's.

Rewriting a page for citation can cost it retrieval

The single most operationally relevant result in the survey is in Section 7.4, and it is a warning rather than a tactic. SAGEO Arena, credited to Kim et al. 2026, matters to the review because it reinstates retrieval and reranking instead of handing the model a fixed context. Across 171,003 documents and 2,700 queries, body-only optimization reduced average top-20 presence by approximately 9 percent, top-10 presence after reranking by 16 percent, and final citation by 6 percent. The survey adds that applying one named optimization system to the body alone can produce larger losses.

The explanation the paper gives is compositional. Fixed-context benchmarks estimate a direct effect on generation. An end-to-end arena estimates the composition of several effects, so a rewrite may perform well once injected while making the document less retrievable or less competitive upstream. If the conditional citation probability rises while the retrieval probability falls, the total effect can be negative. Any operational claim, it says, must specify which of these effects it measures.

Section 7.3 points the same way from a different corpus. C-SEO Bench, credited to Puerto et al. 2025, covers two tasks, six domains, approximately 1,900 queries and 16,360 documents, and finds only three of 54 method and domain combinations significantly positive in the main experiment, with none positive in question answering. Several transformations reduce rank outright, and gains decline as adoption increases toward what the paper calls congested dynamics approaching a zero-sum game. E-GEO, credited to Bagga et al. 2025, reaches a compatible conclusion in e-commerce, where ten of fifteen initial heuristics are neutral or negative.

None of that says writing well is pointless, and the survey does not say so either: Table 4 still rates query to document relevance as strong in controlled settings and extractable evidence as moderate to strong. What it says is that a rewrite tuned for the citation stage is not free at the retrieval stage, which is an argument for fixing what a crawler can actually fetch before optimising the prose it finds. That ordering is the whole reason we measure prose parity first, and what it looks like on real sites is in our capture of six storefronts. Platform pages such as how citation works on Perplexity describe the downstream stage this result cautions about.

  • Average top-20 presence, reduction 9% The survey's wording is approximately 9 percent.
  • Top-10 presence after reranking, reduction 16% The largest of the three reported drops.
  • Final citation, reduction 6% The stage the rewrite was aimed at.
SAGEO Arena, credited in Section 7.4 of arXiv:2607.14035 to Kim et al. 2026, across 171,003 documents and 2,700 queries. Percentage reduction from body-only optimization, as stated in the survey. Reported figures, not a Lantad measurement.

Which stage a scan like ours can actually reach

This is the part of the post where a vendor is supposed to explain that its product solves the problem the paper describes. It does not. Lantad measures one stage of the seven, and the honest thing to do with a paper this clear about scope is to be equally clear about ours.

What an external scan can establish is whether a named crawler is permitted by robots.txt, what the server returns to a fetch that executes no JavaScript, how much of the rendered text survives in that response, and what machine-readable structure is present. That is stage two, crawling and indexing, plus some of the page properties that feed later stages. It is measurable from outside because it is a property of your server's response rather than of a model's internal state.

What no external scan can establish is everything the survey grades Low or Very low. Whether a given engine retrieved your page for a given query. Whether it ranked it inside the context window. Whether the citation you saw yesterday will appear today, which the survey's own reliability section addresses by reporting daily source-level Jaccard overlap of roughly 0.34 to 0.42 across four engines and 45 days, credited to Schulte et al. 2026, with a suggested starting point of seven to eight repetitions per prompt. Whether any of it produced a click. We do not have those numbers for your site and neither does anyone selling you a single score for them.

That is why this site refuses to print a grade when the measurement underneath it failed, which is argued at length in why we withhold a grade, and why the crawler set we do model is published in full at our public crawler directory rather than described. It is also why the identity question matters: a request that says it is GPTBot has only made a claim, as the post on user agents as claims sets out. A paper that grades its own evidence honestly deserves a vendor response that does the same.

Sample Illustrative, not a measurement of any real site.

  • Activation Not observable Whether the engine searched at all is internal to the engine.
  • Crawling and indexing Observable robots.txt per token, server response, text present without JavaScript.
  • Retrieval Not observable The retrieved set is not exposed to the site owner.
  • Reranking and context Not observable Ranking scores are internal.
  • Generation and citation Samplable, not measurable Can be sampled by prompting; the survey reports high run-to-run variance.
  • Absorption and fidelity Not observable Requires counterfactual runs with and without the source.
  • Attention, click, conversion Your analytics, not a scan Graded Very low for causal attribution in Table 5.
Mapping the survey's seven stages against what an external scan of a site can and cannot observe. Illustrative of scope, not a measurement of any site.

What to check on your own site after reading this

The practical reading of the survey is an ordering, not a checklist of tactics. Stage two gates the rest, stage two is the stage almost nobody has studied, and stage two happens to be the one you fully control. Three things are worth confirming before anything downstream.

First, that your robots.txt says what you think it says for each crawler separately rather than in aggregate. Rules are matched per product token, so a file can allow one vendor's search fetcher while blocking its training crawler, and the outcome differs by token. The standard that governs this is RFC 9309, and its handling of an unreachable file is genuinely counterintuitive: a 404 and a 503 point in opposite directions, which we set out in the post on 404 and 503. You can check a file against a specific token with our robots.txt tester.

Second, that the text you want read is in the server response. This is the failure that produces a page which looks complete in a browser and nearly empty to a fetch with no JavaScript engine, and it is invisible from the browser you built the site in. It is the most common structural problem we see, and it is stack-shaped rather than content-shaped: the fixes for a client-rendered app are specific, which is why the guidance is split by framework, starting with the Next.js guide.

Third, that you know which claims about your visibility you can actually verify and which you are taking on trust. That is the survey's real contribution to a buyer rather than a researcher. When a tool reports that a change lifted your citation rate, Table 5 says the causal step from an intervention to durable cross-platform discoverability is graded Low and the step from citation to revenue is graded Very low. Neither grade means the tool is lying. Both mean the burden of proof for that claim has not been met in the literature as of 14 July 2026, and a vendor asserting otherwise is asserting past the evidence rather than from it.

Sample Illustrative, not a measurement of any real site.

Checking stage two before stage five

  • GET /robots.txt, evaluated for one product token at a time allow or disallow, per token
  • GET /page, no JavaScript executed, server response only status and bytes returned
  • Compare text in that response against text in the rendered page share of prose present
  • Parse machine-readable structure present in the response types found or none
  • Only then: sample what an engine says, repeatedly, across days a distribution, not a number
The order the survey's staging implies for a site owner: confirm access and readability before optimising for citation. Illustrative sequence, not a capture of any real site.

Related

Common questions

Does the survey say generative engine optimization does not work?

No. It says the evidence is conditional on retrieval. Its Table 5 grades as High confidence that a document already placed in the context can causally alter its rank, citation or use, and grades as Low the claim that a white-hat intervention durably improves organic discoverability across multiple engines. Its conclusion states that within the reviewed corpus no technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream clicks and conversions.

Where does the 40 percent figure for GEO come from?

From the foundational GEO paper, and the survey rejects it as a general claim. Table 5 lists GEO increases visibility by 40 percent among claims rejected as general, describing the figure as a relative maximum on one metric under a specific configuration. The introduction adds that it describes a relative visibility gain in a simulator in which five documents have already been placed in context.

Why does the crawling stage matter more than its share of the research?

Because it is second in the chain. Figure 1 of the survey orders the pipeline as activation, crawling and indexing, retrieval, reranking and context, generation and citation, absorption and fidelity, then attention and clicks, and its caption says far fewer studies observe crawling, organic retrieval or user behavior. A page that is not fetched or not readable without JavaScript cannot reach the stages that are heavily studied.

Did Lantad run any of the measurements in this post?

No. Every figure here is transcribed from arXiv:2607.14035, posted 15 July 2026, including the results it credits to other papers such as SAGEO Arena and C-SEO Bench. Lantad measures crawler access, what a page returns without JavaScript, prose parity and machine-readable structure, which is the survey's second stage. It does not measure retrieval, citation, absorption or conversions.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.