# AI visibility took seven runs per prompt to settle

> A study from the University of St. Gallen posted to arXiv on 8 April 2026 queried four AI search engines with 32 prompts over 45 to 46 days, then repeated each prompt up to 10 times inside one day. A per-brand visibility rate estimated from a single run carried a standard error of 0.370, and seven runs were needed to bring it under 0.10.

- Canonical page: https://lantad.co/blog/ai-visibility-took-seven-runs-per-prompt-to-settle
- This file: https://lantad.co/blog/ai-visibility-took-seven-runs-per-prompt-to-settle.md
- Last substantive update: 2026-08-16

## Key facts

- **Published:** 2026-08-16
- **Category:** Findings
- **Author:** Lantad
- **Length:** 4054 words
- **Takeaway 1:** A study by Julius Schulte, Malte Bleeker and Philipp Kaufmann of the University of St. Gallen, posted to arXiv on 8 April 2026 as arXiv:2604.07585, queried ChatGPT, Gemini, Google AI Mode and Perplexity with eight prompts in each of four verticals across a 45 to 46 day window running 24 January to 20 March 2026.
- **Takeaway 2:** Cited sources overlapped between two consecutive days with a mean Jaccard similarity of 0.336 to 0.423 depending on the vertical, which the authors read as roughly 65 percent of the cited sources changing from one day to the next, across 4,044 consecutive-day pairs.
- **Takeaway 3:** Repeating the same prompt up to 10 times inside a single 24 hour window produced pairwise source overlap of 0.321 to 0.434 across 3,409 comparisons, the same range as the day to day figures, so most of the instability came from the model's own sampling rather than from index drift.
- **Takeaway 4:** A subsampling analysis of that 10 run dataset put the standard error of a per-brand detection rate estimated from one run at 0.370, falling to 0.081 at seven runs and 0.062 at eight, measured across 1,216 per-brand series.
- **Takeaway 5:** Lantad measured none of this. Lantad's own prompt runs ask each prompt once per engine per run and reschedule at most twice a week, which is a setting recorded in core/src/config.ts rather than a finding about how many runs are enough.

## Summary

There are two kinds of number on this site and they do not behave the same way. One kind asks whether a crawler can fetch a URL, whether robots.txt permits the fetch, and whether the words a reader sees are present in the HTML that arrives. Ask that twice and you get the same answer twice, because nothing in the chain is sampling from a distribution. The other kind asks what a model says about you when somebody types a question, and that number is a draw.

A paper posted to arXiv on 8 April 2026 puts a size on how big a draw it is. Three researchers at the University of St. Gallen ran a repeated-measurement study of [AI visibility](https://lantad.co/glossary/ai-visibility) across four engines and four product categories, first daily for six weeks, then up to ten times in a single day to separate the model's own randomness from anything changing in the world. The headline is not that answers vary, which everybody in this market already suspects. It is that the variation is large enough to be measured, that it survives holding the day constant, and that the authors can say how many repetitions it takes before a visibility figure means anything. Lantad did not run this study and has measured none of it: every figure below is read from [the paper itself](https://arxiv.org/abs/2604.07585) on 16 August 2026.

## What the study measured, and on which engines

The design is worth stating precisely, because the value of the result rests entirely on the repetition. The authors built two datasets. The first records the daily results of four AI search engines across four Swiss-German campaign verticals over a 45 to 46 day window running 24 January to 20 March 2026. The second re-issues the same prompts several times on the same calendar day, which is what lets the paper separate a model that is simply random from a web that is simply changing.

The engines are ChatGPT, Gemini, Google AI Mode and Perplexity. Google AI Overviews was collected too and then excluded from every analysis, on the stated grounds that it generates summary snippets inside ordinary search results rather than operating as a dedicated AI search interface, and so differs in interaction mode and citation behaviour from the other four. That exclusion is worth carrying forward if you are tempted to read these figures as covering everything Google does, because [AI Overviews is a different surface](https://lantad.co/how-to-get-cited/google-ai-overviews) with different mechanics.

The prompts were derived from high search volume SEO keywords, entered into Google, and expanded through the People Also Ask feature into conversational questions. Eight prompts were selected per campaign across four verticals: Telecommunications, Real Estate Sales, Sporting Goods and Consumer Electronics. That is 32 prompts in total, which is small, and the paper does not pretend otherwise. Brand detection ran against a campaign-specific lexicon of 32 to 51 canonical brands per vertical.

Coverage was not uniform. Gemini had sporadic gaps in the period, answering on 22 to 26 days out of 45 or 46 while the other three engines answered on 38 to 44. One day, 30 January 2026, was excluded outright because its citation volume ran at roughly twice the daily average. That is ordinary field collection against consumer products nobody promised would stay up, and it is the sort of thing a [generative engine optimization](https://lantad.co/glossary/geo) vendor quoting one clean percentage will not usually mention. The paper puts it in a table.

Two similarity metrics do the work throughout. Jaccard similarity measures plain set overlap between two observations, the size of the intersection over the size of the union, and it ignores order. Rank Biased Overlap at p equal to 0.9 weights items near the top of the list more heavily. Reading both together distinguishes "the same sources came back in a different order" from "different sources came back", and the paper reports both for every cut. If you are tracking whether [ChatGPT cites you](https://lantad.co/how-to-get-cited/chatgpt) or whether [Perplexity does](https://lantad.co/how-to-get-cited/perplexity), those are two different failure modes and only the second one loses you the citation.

## Two thirds of the cited sources changed from one day to the next

Across all four campaigns over the six week window, the day-to-day Jaccard similarity for cited sources averages between 0.336 and 0.423, aggregated over 4,044 consecutive-day pairs. The authors spell out what that means in plain terms: a Jaccard value of 0.35 implies that on average only about 35 percent of the cited sources overlap between two consecutive days, so roughly 65 percent of all sources change overnight. Consumer Electronics was the least stable at 0.336 and Telecommunications the most stable at 0.423, and the standard deviations sit between 0.243 and 0.293, which is to say the spread around those means is nearly as large as the means.

Rank Biased Overlap is consistently lower, at 0.206 to 0.256. That is the more uncomfortable pair of numbers. It says the source sets are not merely being reshuffled while the same domains stay in play: the ordering moves too, and it moves more than the membership does.

Brand mentions were more stable than source citations, but not by as much as a marketer would want. Restricted to the three campaigns whose mean brand-detection rate cleared 70 percent, the day-to-day brand Jaccard runs 0.453 to 0.589 over 2,924 consecutive-day pairs, with RBO between 0.187 and 0.304. Real Estate Sales was excluded from brand analysis entirely at a 53.6 percent detection rate, which the authors attribute to generic tax and investment queries that the models answer without naming any specific company. Sporting Goods showed the lowest brand Jaccard at 0.453, and the paper's explanation is the obvious one: when a large pool of running shoe brands is substitutable, the model draws from a wide set across days.

There is a concentration finding sitting alongside the instability finding, and the two of them together describe the shape of the problem. The mean Gini coefficient of source citations across all campaigns and engines is 0.715, with Google AI Mode highest at 0.782 and Perplexity lowest at 0.671. A small number of domains take most of the citations, and which small number it is keeps changing. That is consistent with what other work has reported about where AI citations land: a separate preprint we covered found that [a brand's own domain drew 2.9 percent of AI citations](https://lantad.co/blog/ai-citations-mostly-point-at-other-companies), and a controlled study found that [four content factors decided the first citation](https://lantad.co/blog/four-factors-decided-the-first-citation) while presentation and formatting did not. None of that helps if the measurement of your own position is drawn from a distribution this wide.

For anyone working on [answer engine optimization](https://lantad.co/glossary/aeo), the practical consequence is about attribution rather than tactics. If you change something on your site and your mention rate moves from 0.4 to 0.5 between two single-run checks a week apart, the study says you have learned essentially nothing about your change. The same swing appears between two runs where nothing was changed at all. Interventions that genuinely did move the needle, such as the case where [a curated trusted domain list moved citations from 12 to 21 percent](https://lantad.co/blog/trusted-domain-list-moved-citations-12-to-21-percent), were measured across hundreds of answers rather than across a before and an after.

## Repeating the same prompt within a day was almost as unstable

Everything above could in principle be the web moving rather than the model wobbling. Indexes refresh, pages get published, ranking systems update. The second dataset exists to rule that out, and it is the part of the paper that changes how a visibility number should be read.

The authors issued eight prompts per campaign up to ten times in succession to all four engines, then compared only pairs whose timestamps fall within 24 hours of each other. A second filter kept only runs returning at least one extracted citation, since a zero-citation response would otherwise inflate the overlap score with a false match on emptiness. That filter passed 75.4 percent of runs overall, and ChatGPT was lowest at 42.2 percent, which the paper attributes to its suppressing web search on definitional queries. After both filters, 3,409 pairwise source comparisons remained.

Under classical search assumptions you would expect near-identical results for the same query issued minutes apart. The actual pairwise Jaccard for sources averages 0.321 to 0.434 across campaigns. That is the same range as the day-to-day figures. Holding the calendar day fixed removed almost none of the instability, which is the paper's central result: most of what looks like the web changing under you is the generation process sampling differently.

Brand mentions behaved the same way, at 0.327 to 0.477 across 1,027 to 1,235 pairs per campaign, with within-campaign standard deviations around 0.30. Some prompts returned nearly identical brand sets across all ten runs while others changed almost entirely, and the authors note the pattern: specific product queries attract more consistent sources and brands than broad generic ones. Query specificity, not just engine behaviour, drives how repeatable an answer is.

The per-engine breakdown is the most quotable table in the paper and the one most likely to be quoted wrongly, so the caveat travels with it. Within 24 hours, ChatGPT's source overlap averaged 0.233 and Perplexity's 0.282, while Gemini reached 0.505. Those source columns include only runs with at least one extracted citation, so they describe stability among citing runs and say nothing about how often each engine cites at all. Brand overlap ran the other way round, with Perplexity highest at 0.492 and Google AI Mode lowest at 0.375. Whether you are working on [getting cited by Claude](https://lantad.co/how-to-get-cited/claude) or by any of the four here, the engine you are least able to move may simply be the one whose answers are noisiest, and the paper's own conclusion is that these are distributions rather than positions. The authors put it as an inclusion and exclusion dynamic: unlike classical search, where a page slips down the ranking, a brand here is either in the answer or entirely absent from it, and [the full text](https://arxiv.org/abs/2604.07585) makes that contrast the reason snapshot metrics mislead. It is also why [entity confidence](https://lantad.co/glossary/entity-confidence) and citation presence are worth tracking as separate things.

## How many runs before the number stops moving

The most directly useful part of the paper is an appendix. Having collected ten runs per engine and prompt, the authors ask what a smaller number of runs would have told them. For each group and each subsample size from one to nine they draw 2,000 random subsamples without replacement and record the mean detection indicator for each individual brand, treating the ten run mean as the best available proxy for the true detection probability. The standard deviation across those 2,000 subsample means is the estimated standard error of an n run estimate. Doing this per brand rather than collapsing to a campaign-level "any brand" indicator yields 1,216 per-brand series across the three qualifying campaigns.

The curve is steep and then flat. A single run gives a standard error of 0.370, which the authors describe as essentially uninformative: a true detection rate of 50 percent could appear anywhere across a nominal 95 percent interval running from below zero to above one, clipped to the unit interval in practice. Five runs bring it to 0.123. Seven runs bring it to 0.081, the first point under 0.10, with a 95 percent confidence half-width of 0.158. Eight runs give 0.062 and a half-width of 0.121. Source coverage converges more slowly, needing eight runs to reach a standard error of 0.096, because which specific URLs appear is more variable than whether a brand is named at all.

The authors are careful about what those figures are worth, and the caveat should travel with the numbers. Sampling without replacement from a pool of ten introduces a finite population correction, so the standard error at nine runs is mechanically smaller than it would be with truly independent fresh runs. The reported values are a lower bound. They also note that seven runs is adequate for detecting large differences, such as a brand appearing in 80 percent of runs against one appearing in 20 percent, and insufficient for fine-grained ranking of brands with similar visibility. A vendor dashboard that ranks you eleventh against a competitor in tenth place is making a claim this evidence cannot support.

This is where the finding stops being about somebody else's product and starts being about ours. Lantad's prompt tracking asks each prompt once per engine within a run: the worker loops over the selected prompts and appends a single answer per engine per prompt. Repetition comes from the schedule instead, and the schedule is a plan setting rather than a statistical choice. The plan rows in core/src/config.ts set scheduled reruns at one per week on the lower tiers and two per week at the top, and the free preview runs once with no scheduled rerun at all. Measured against this paper's threshold, a single [prompt run](https://lantad.co/prompts) is one sample, and a week of runs at the highest cadence is two. Those numbers are design decisions about model spend, not findings about sufficiency, and they are visible on the [pricing page](https://lantad.co/pricing) as what they are. Saying so is cheaper than the alternative, which is letting a customer read a one-run mention rate as a measurement. Our [methodology page](https://lantad.co/methodology) already commits to naming what was not measured, and run count belongs on that list.

## What the authors say the data cannot support

A paper that lists its own limitations at this length is doing the reader a favour, and repeating them is part of citing it honestly.

All data were collected from servers located in Switzerland, so every prompt was served with Swiss IP addresses and locale settings, and the prompts themselves were written in German. The authors state plainly that geo-personalised index selection, language weighting and citation patterns may all differ elsewhere, and that results may not generalise to other regional or linguistic markets. Nobody should read 0.336 as a number about the English-language web.

The ChatGPT data contained an OpenAI image delivery CDN, images.openai.com, appearing as a spurious cited domain in 889 cases, which the paper puts at 5.8 percent of its ChatGPT citations for that period. It is filtered from every calculation. The same limitations section is the reason the analysis is restricted to the January to March window at all: collection method consistency across the four engines could not be maintained outside it, and the authors recommend that future studies pin collection method and model version across the full observation period.

Brand detection relies on substring matching against a fixed lexicon, so brands referred to by synonym, abbreviation or paraphrase are missed, and generic words that happen to be substrings of brand names produce false positives. The 70 percent detection-rate threshold mitigates the worst of it without eliminating it. And 57.8 percent of ChatGPT runs carried zero citations because the product only activates web search for certain queries, which is a large hole in any attempt to reason about how often that engine cites anything.

One inconsistency is worth flagging because we are asking readers to trust these figures. The convergence appendix states in one sentence that the method yields 1,216 per-brand series, and in the next that the mean standard error is reported across 1,259 series. The table note, the method description and the figure caption all say 1,216, so that is the figure used above, but the two do not agree in the published text and we have not resolved which is correct. Separately, the appendix on how long an observation window needs to be is present in the table of contents with headings but carries no results in the version read on 16 August 2026, so no window figure appears in this post.

That list is longer than most vendor statistics come with, and it is the reason this study is citable at all. It is the same standard applied to [our own crawlability research](https://lantad.co/research), where the [sample and its composition](https://lantad.co/research/crawlability-study) are published alongside the headline number. It is also worth setting against what platforms themselves disclose: Google's own generative AI report in Search Console [counts impressions and nothing else](https://lantad.co/blog/generative-ai-report-counts-impressions-only), with no clicks and no click-through rate, so the first-party alternative to a noisy third-party estimate is currently a narrower number rather than a firmer one.

## Which of your AI visibility numbers survive being taken once

The useful move after reading this paper is not to distrust every AI visibility figure. It is to sort your figures into the ones that are draws from a distribution and the ones that are not, because they need completely different treatment and most dashboards mix them.

A mention rate, a share of voice, a citation count, a sentiment score: every one of those is produced by asking a model something and reading what comes back, so all of them inherit the variance measured above and none should be reported from a single observation. The paper's recommendation is at least seven runs per prompt per day for brand visibility and at least eight where source-level coverage matters, and its broader point is to treat visibility as a probability of appearing rather than a position in a list. If your tool shows a single number with no interval and no run count, ask which of the two it is.

The other half of the ledger behaves differently, and this is what Lantad's scanner actually does. Whether [an AI crawler](https://lantad.co/glossary/ai-crawler) receives your content is not a sampling problem. A request either returns 200 or it does not. A robots.txt group either matches a product token or it does not, and [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html) specifies that matching precisely enough that two implementations should agree. The prose either exists in the delivered HTML or it arrives only after JavaScript runs. Run [a check of what GPTBot sees](https://lantad.co/tools/what-gptbot-sees) twice and the answer is the same twice, barring an actual change to the site or an outage. That reproducibility is why a readability check can carry a grade while an answer-tracking figure should carry an interval.

Reproducibility is a property of the measurement, not a guarantee that the summary built on top of it is well designed. We published that [a robots.txt blocking every citation-capable crawler still graded B](https://lantad.co/blog/a-robots-txt-that-blocks-every-citation-crawler-still-grades-b) under our own weights, which was a reporting fault in a perfectly repeatable check. What determinism buys is narrower and still worth having: a disagreement between two runs is a real event to investigate rather than the expected behaviour of the instrument.

The honest reading of this paper for a site owner is a division of labour. Fix the deterministic layer first, because it is cheap to verify, it stays fixed, and a page a crawler cannot read cannot be cited by any engine at any sample size. Then treat the probabilistic layer as sampling: more runs, more prompts, intervals rather than points, and scepticism toward any week-on-week movement that a single observation produced. When you look at [what AI says about you](https://lantad.co/what-ai-says), read it as one draw from the distribution this study measured.

## Questions and answers

**How many times should the same prompt be run before an AI visibility figure is reliable?**

The study behind this post recommends at least seven runs per prompt per day for brand-level visibility, the first point where the standard error of a per-brand detection rate falls below 0.10, and at least eight runs where source-level coverage matters. It also states that seven runs is adequate for detecting large differences and insufficient for finely ranking brands whose visibility is similar.

**Is AI answer variation caused by the web changing or by the model?**

Mostly by the model. When the authors repeated the same prompt up to ten times inside a single 24 hour window, source overlap averaged 0.321 to 0.434 across campaigns, which is the same range as the day-to-day figures over six weeks. Holding the calendar day fixed removed almost none of the instability.

**Does this mean AI crawler readability checks are unreliable too?**

No, and the difference matters. Whether a crawler receives a 200, whether a robots.txt group matches a product token under RFC 9309, and whether prose is present in the delivered HTML are deterministic outcomes that repeat. The variance measured in this study applies to what a model says, not to whether your page can be fetched and parsed.

**How many times does Lantad ask each prompt?**

Once per engine per run. Repetition comes from the schedule: plan settings in core/src/config.ts provide one scheduled rerun per week on the lower tiers and two per week at the top, and the free preview runs once with no scheduled rerun. Those are spend decisions rather than statistical ones, and by this study's threshold a single run is a single sample.

---

Lantad measures whether AI crawlers can actually read a page: it fetches as a non-rendering
crawler, renders as a browser, and reports the gap. Free scan, one URL, no signup.

Method and weights: https://lantad.co/methodology | All pages as markdown: https://lantad.co/md | Crawler policy: https://lantad.co/bot
