Blog / How many prompts an AI visibility measurement needs

How many prompts an AI visibility measurement needs

A study of 374,052 citations across Perplexity, OpenAI SearchGPT and Google Gemini found repeated runs of the same query share as little as 0.29 of their cited domains, and that a confidence interval five percentage points wide on citation share needs roughly 40 queries on one platform and 150 or more on another.

In short

  • Quantifying Uncertainty in AI Visibility, by Ronald Sielinski, was posted to arXiv as 2603.08924 on 9 March 2026 and revised on 9 June 2026. It submitted the same 200 queries per topic to Perplexity Search, OpenAI SearchGPT and Google Gemini once a day for nine days, and again every ten minutes for about four hours.
  • Repeated runs of the same query returned different sources: the median domain level Jaccard overlap was 0.29 to 0.31 on Gemini, 0.33 to 0.40 on SearchGPT and 0.50 on Perplexity, across a daily dataset of 374,052 citations.
  • The paper's sample size guidance for a 95 percent confidence interval five percentage points wide on citation share is roughly 40 to 50 queries on Gemini, about 90 to 100 on Perplexity and 150 or more on SearchGPT.
  • A control that stored a SHA-256 checksum of every cited page found the overwhelming majority of transitions unchanged between samples, so the citation variability the paper measures is platform behaviour rather than publishers editing pages.
  • Lantad's own answer tracking asks each engine once per prompt per run, and the plan settings in core/src/config.ts carry 5, 25 or 50 prompts per run, so a single run is one draw and is reported with an engine name and a timestamp rather than as a stable share.

An AI visibility figure is almost always printed once. A tool sends a set of questions to an answer engine, counts how often a domain is cited or a brand is named, and reports a percentage. That percentage then gets treated as a property of the site, in the way a response time or a link count is a property of a site. A paper posted to arXiv on 9 March 2026 and revised on 9 June 2026 asked what happens when you send the same questions again ten minutes later, and found that a large part of an AI visibility measurement belongs to the measurement rather than to the site being measured.

This post reports what that study did and what it found, in its own figures, and then does the part that matters when buying AI visibility measurement: separating the numbers on a report that are repeatable from the numbers that are one draw from a distribution. The two look identical on a dashboard and they support completely different decisions.

Lantad did not run this study and has not replicated it. Nothing below is a Lantad measurement, and every figure comes from the paper with its sample size attached. What this site can add is the boundary. A scan of your own pages is repeatable to the byte, an answer engine's citation set is not, and the methodology behind any AI visibility number ought to make clear which of the two you are reading.

  • Platforms sampled 3 Perplexity Search, OpenAI SearchGPT and Google Gemini, each given the same 200 queries per topic across three consumer product topics.
  • Sampling regimes 9 daily, 25 rapid Once a day over nine consecutive days, plus one topic sampled every ten minutes for about four hours.
  • Median citation overlap 0.29 to 0.50 Domain level Jaccard between repeated runs of the same query. Lowest on Gemini, highest on Perplexity.
  • Queries for a five point interval 40 to 150 plus Sample size the paper reports for a 95 percent confidence interval spanning five percentage points on citation share, by platform.
Design and headline results of Quantifying Uncertainty in AI Visibility, Ronald Sielinski, arXiv 2603.08924v2, dated 9 June 2026. Reported from the paper, not measured by Lantad.

What the study asked, and how many times it asked it

The paper is Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement by Ronald Sielinski, whose affiliation line reads IQRush. Version 1 was submitted on 9 March 2026 and version 2, the one read for this post, is dated 9 June 2026, filed under stat.AP and published under CC BY 4.0. Its own keyword line names three fields: AI visibility, answer engine optimization, and generative engine optimization.

Three topics were chosen for different market structures. Bird feeders is described as a fragmented market with no dominant brand, running gear as brand concentrated and multi product, and multivitamins for adults as a health and life sciences context with institutional sources alongside commercial ones. For each topic a model generated 200 queries conditioned on ten query types, among them value for money, feature comparisons, common complaints and expert recommendations. Uniqueness was not enforced, on the stated ground that frequently generated queries are likely to be frequently asked ones, and the repetition rate ran from 7.5 percent on running gear to 24.0 percent on multivitamins.

The same 200 queries per topic went to all three platforms under two regimes. The daily regime submitted them once a day for nine consecutive days, dated 3 to 11 February 2026 in the figure captions. The high frequency regime took running gear alone and submitted it every ten minutes for roughly four hours, yielding 25 samples per platform, chosen so that the underlying pages had almost no opportunity to change. After extraction and deduplication the daily dataset held 374,052 citations and the high frequency dataset held 379,276.

One result arrives before any statistics. Mean citations per response were 39.6 to 43.1 on Gemini, 20.0 to 22.5 on Perplexity and 5.9 to 7.1 on SearchGPT. Raw citation counts cannot be compared across those platforms at all, which is the same category error as treating a grounded model answer as an AI Overview, the subject of a Gemini citation tool is not an AI Overview tool. The unit of output differs per engine before anybody computes anything with it.

  • Gemini, high frequency 47.1 25 samples, ten minute intervals
  • Gemini, daily 43.1 9 samples, median 40
  • Perplexity, high frequency 22.7 25 samples, median 22
  • Perplexity, daily 22.5 9 samples, median 22
  • SearchGPT, high frequency 6.7 25 samples, median 5
  • SearchGPT, daily 6.6 9 samples, median 5
Mean citations per response, from Tables 2 and 3 of the paper, running gear topic on all three platforms under both sampling regimes. Bars are drawn on a 0 to 100 scale, so each bar is the mean count against a 100 citation ceiling.

The same query returns a different set of sources

Before aggregating anything, the paper checks whether repeated runs of one query return the same sources. It compares citation sets at domain level, which is the generous test: two responses match on a domain if both cite at least one URL from it, and the paper notes that URL level matching would be strictly more variable.

Median Jaccard overlap across repeated runs of the same query was 0.29 to 0.31 on Gemini, 0.33 to 0.40 on SearchGPT and 0.50 on Perplexity. Read plainly, two runs of one question on Gemini agree on roughly three cited domains in ten, and the steadiest of the three agrees on half. The identical citation rate, where two responses to the same query share every source, runs from 0.01 to 0.10 percent on Gemini up to 3 to 8 percent on the other two. The zero overlap rate, where two runs of one query share nothing at all, is 6 to 9 percent on SearchGPT and 1 to 2 percent on Perplexity.

An obvious explanation would be response length: more citations, more room to differ. The paper tests that and rejects it. Jaccard similarity stays essentially flat across the observed range, near 0.30 on Gemini across responses carrying roughly 10 to 70 citations, near 0.40 to 0.42 on SearchGPT across 3 to 12, and near 0.50 on Perplexity across 5 to 35. What determines overlap is which platform you asked, not how much it said. The paper reads this as evidence that variability comes from platform level retrieval and ranking rather than from the combinatorics of larger citation sets.

The consequence for anyone reading a single citation table is direct. If your domain appears in one run and not the next, neither run is wrong, and neither is evidence that anything changed on your site. It is also why a report on how ChatGPT surfaces sources is more useful when it names the engine and the retrieval mode than when it prints one number, and why a brand named once is a weaker claim than the same brand named in most of a set.

  • Perplexity, bird feeders 0.50 198 queries, 7,128 pairs
  • Perplexity, multivitamins 0.50 199 queries, 7,164 pairs
  • Perplexity, running gear 0.50 198 queries, 7,128 pairs
  • SearchGPT, bird feeders 0.40 180 queries, 6,480 pairs
  • SearchGPT, multivitamins 0.38 180 queries, 6,480 pairs
  • SearchGPT, running gear 0.33 169 queries, 6,084 pairs
  • Gemini, bird feeders 0.31 189 queries, 6,804 pairs
  • Gemini, multivitamins 0.29 192 queries, 6,912 pairs
  • Gemini, running gear 0.29 192 queries, 6,912 pairs
Median domain level Jaccard overlap between repeated runs of the same query, from Table 4 of the paper, all nine platform and topic combinations. Values run 0 to 1 and bars are drawn at 100 times the value, so a full bar would be perfect agreement.

Where the noise floor sits in an AI visibility measurement

Section 5.6 of the paper is the one to read if you have ever compared two domains on a dashboard. It takes a single sample of responses, resamples it with replacement 1,000 times, and recomputes citation share on each replicate, which produces an empirical confidence interval around a number that is normally printed alone.

The intervals are wide. On SearchGPT the span for most frequently cited domains runs 3 to 6 percentage points. On the other two platforms they are narrower and still large next to the differences people act on: everydayhealth.com on multivitamins on Gemini has a citation share of 6.0 percent with an interval spanning 3.2 percentage points, and runnersworld.com on running gear on Perplexity has 13.4 percent with a span of 4.7. The paper's own summary is that many apparent differences between domains "fall within the noise floor of the measurement process".

One case shows what a single run can do unaided. On Gemini bird feeders, nationalgeographic.com registered a citation share of 0.032 on the first day with a bootstrap interval of 0.024 to 0.042, against a cross sample mean of 0.005 over all nine days. Somebody reading day one alone would have called it a top cited domain, and every later sample contradicts that. On SearchGPT multivitamins, yahoo.com ranged from about 0.092 to 0.160 inside the nine day window, a factor of nearly two for the top ranked domain.

Rank order moves too, and not only at the top. The paper's distribution wide analysis compares consecutive daily samples with a share weighted Spearman correlation and counts a pair as stable at 0.9 or above. Gemini running gear was the cleanest result in the dataset, with seven of eight pairs stable and a mean of 0.913, while Gemini multivitamins managed two of eight at a mean of 0.806. SearchGPT multivitamins and running gear produced zero pairs precise enough to judge at all. Comparing the first day against the last, Gemini bird feeders and multivitamins fell to 0.689 and 0.726, a cumulative drift invisible in any single adjacent comparison.

This is the arithmetic behind a rule this site already applies to itself: we withhold a grade rather than print a number the measurement did not support, and the research page publishes a sample size beside every figure. A confident percentage with an undisclosed interval of plus or minus three points is not a better answer than an honest range. It is the same answer with the uncertainty deleted.

  • everydayhealth.com 6.0 percent, span 3.2 pts Multivitamins on Gemini. The interval is over half the size of the estimate it surrounds.
  • runnersworld.com 13.4 percent, span 4.7 pts Running gear on Perplexity. The largest shares carry the largest absolute intervals.
  • nationalgeographic.com 0.032 on day one Bird feeders on Gemini. Cross sample mean 0.005, and the day one interval of 0.024 to 0.042 misses it entirely.
  • SearchGPT rank stability 0 of 8 pairs stable Bird feeders had five pairs precise enough to judge and none stable. Multivitamins and running gear had none precise enough to judge.
Four worked examples from Sections 5.6 and 5.8 of the paper, quoted with the platform and topic each was measured on. Bootstrap intervals use 1,000 replicates at 95 percent.

How many queries a five point confidence interval needs

The most directly useful part of the paper is Section 5.7, which plots confidence interval width against the number of queries in the sample and marks where each platform crosses a target width. The targets are stated as practical benchmarks rather than derived quantities: 0.05 for citation share and 0.15 for citation prevalence, where prevalence counts how broadly a domain appears across responses rather than how many citations it collects.

For citation share, Gemini converges fastest, crossing 0.05 at roughly 30 queries on bird feeders and 40 to 50 on the other two topics. Perplexity crosses at roughly 90 to 100 across topics. SearchGPT needs 150 or more, and its curve is not monotonic: the width narrows, widens again, and only smooths after about 120 queries, which the paper partly attributes to subsamples being drawn from an increasingly exhausted pool as the count approaches the full 200 rather than to any real stability.

For citation prevalence the ordering reverses. SearchGPT crosses the 0.15 target earliest at 60 to 80 queries, Perplexity at 100 to 140, and Gemini needs 140 to 150. Prevalence is easier to estimate on the platform that emits fewest citations per response. Two metrics that both get called visibility can therefore have opposite sample size requirements on the same platform, which is worth knowing before setting one vendor's number against another's, as our own comparison pages do.

The non monotonic curves also rule out the obvious shortcut. You cannot collect queries until the interval looks narrow and then stop, because the width can narrow by luck and widen again, and the paper's advice is to fix the sample size in advance from prior measurements of that platform and topic. Anyone sizing a generative engine optimization programme around a prompt set should read these numbers as a floor for one metric, on one platform, on one topic. The paper declines to publish a general formula and lists deriving one as future work, which is a more honest position than most of the figures in circulation.

  • Gemini, citation share About 30 queries on bird feeders, 40 to 50 on multivitamins and running gear. The decline closely tracks the theoretical curve.
  • Perplexity, citation share Roughly 90 to 100 queries across all three topics, with visible bumps on two of them but all reaching the target inside the 200 query sample.
  • SearchGPT, citation share 150 or more, with non monotonic convergence the paper attributes to the citation distribution shifting across the query sequence.
  • Citation prevalence, all three The order reverses: SearchGPT 60 to 80, Perplexity 100 to 140, Gemini 140 to 150.
Sample size at which the 95 percent confidence interval crosses the paper's target width, from Section 5.7. Share target 0.05, prevalence target 0.15, measured across three topics of 200 queries each.

The variation is not the cited pages changing

A reasonable objection to all of the above is that the web moved. If cited pages were edited between samples, changing citation sets would be a fact about publishers rather than about engines. The paper anticipates that and runs a control.

During every collection job the HTML of cited URLs was scraped, human readable text extracted with Trafilatura, and a SHA-256 checksum computed and stored against the URL and the job identifier. Each domain transition is then classified: unchanged when the hash matches the previous job, changed when at least one URL hash differs, and unknown when hash data is missing on either side. Unknown appears only on the first job of each panel, which is a property of having nothing to compare against rather than a gap in coverage.

The result is that the overwhelming majority of transitions are unchanged, so the citation share volatility sits on top of largely stable source material. The paper is careful about the converse and says so in terms: an identical hash rules out content change, but a differing hash does not prove editorial change. Its example is runnersworld.com on Perplexity, which shows a changed hash on every job after the baseline, consistent with advertisement slots, session tokens or timestamps altering the bytes without altering the text a reader sees.

That distinction is familiar from the scanning side of this problem. Two fetches of one page can differ byte for byte and mean nothing, which is why prose parity is measured on extracted text rather than on raw markup, and why what GPTBot sees reports the text a crawler receives rather than a diff of HTML. The four hour window makes the point on its own: rank crossings inside a window that short cannot be explained by editing, and the paper's conclusion is that the variability is structural rather than content driven.

Sample Illustrative, not a measurement of any real site.

One cited URL, job N against job N minus 1

  • GET cited URL during collection job HTML stored
  • extract human readable text with Trafilatura text only
  • compute SHA-256 over extracted text checksum stored by URL and job
  • hash identical to previous job unchanged: content change ruled out
  • hash differs from previous job changed: necessary, not sufficient
  • no hash on either side unknown: first job only
The content change control described in Sections 4.5 and 6 of the paper, as a sequence for one cited URL across two collection jobs. Illustrative of the procedure, not a measurement of any real page.

What this means for the numbers on a Lantad report

Now the uncomfortable half, because this site sells measurement and part of what it measures is answer engine output.

Lantad's answer tracking asks each engine once per prompt per run. The registry in core/src/engines.ts lists five engines in a fixed order, and the plan settings in core/src/config.ts decide how many prompts a run carries: 5 on the free tier, 25 on the standard plan and 50 on the two larger ones, which the pricing page sells by that breadth. Those counts are decisions somebody made, not findings. Read against this paper, one run of 25 prompts is one draw, and the honest description of a single what AI says result is that this engine said this, once, at this timestamp.

Two differences matter before anyone maps the paper's sample sizes onto ours. It measures citation share and prevalence over a topic, aggregating hundreds of queries into a percentage per domain, whereas Lantad's prompt tracking asks whether a specific page can ground a specific question and whether a name appears in the answer, which is closer to a per question boolean than to a share of a distribution. And its query counts are tied to a target interval width on a share, so they do not transfer to a different metric without new work. The direction transfers even so: run it again and expect the set of rival names to move, because that set is extracted from answer text that is itself a sample.

The other half of a report does not have this problem, and the contrast is the useful part. Whether your robots.txt allows a given AI crawler token, whether your structured data identifies a publisher, whether the text a crawler receives carries the prose a browser shows: those are properties of bytes your own server returns, and two scans an hour apart agree unless you changed something in between. A robots.txt tester returns the same verdict every time it is run against the same file, and that is not a modest claim next to the numbers above. That asymmetry is the reason page level output carries a grade while answer engine output carries an engine name and a date.

Publishing this is the same discipline as publishing that llms.txt has no measured effect while shipping an llms.txt tool, and as reporting that nearly a third of AI Overview citations are not on page one from somebody else's study rather than pretending to have run it. A measurement product that cannot say which of its own numbers are noisy is not measuring, it is decorating.

Repeatable between two scans

  • robots.txt rules evaluated per crawler token
  • Crawler visible text against browser rendered text
  • Structured data present, valid, and naming a publisher
  • HTTP status and redirect chain for the URL
  • Same input, same output, unless the site changed

One draw from a distribution

  • Which sources an engine cites for a question
  • Whether a brand is named in a given answer
  • Which rival names appear alongside yours
  • Rank order inside a citation table
  • Median overlap between two runs: 0.29 to 0.50
Which half of a report is repeatable and which is a sample. The left column describes properties of bytes a server returns; the right describes output the paper shows is non deterministic.

Related

Common questions

How many prompts does an AI visibility measurement need?

For citation share with a 95 percent confidence interval five percentage points wide, the paper reports roughly 40 to 50 queries on Google Gemini, about 90 to 100 on Perplexity Search and 150 or more on OpenAI SearchGPT, measured on three consumer product topics of 200 queries each in February 2026. Those figures apply to that metric on those platforms, and the paper explicitly leaves a general formula to future work.

Why do two AI visibility tools report different numbers for the same site?

Part of the gap is the measurement rather than the tools. Repeated runs of one query share between 0.29 and 0.50 of their cited domains depending on the platform, so two tools sampling at different moments will disagree even with identical prompt sets and identical methods. The rest of the gap comes from genuinely different prompt sets, engines and metrics, which is why a vendor should tell you all three.

If my citation share dropped, did something change on my site?

Not necessarily. The paper stored a SHA-256 checksum of every cited page at every collection and found the overwhelming majority unchanged between samples, and its four hour high frequency window recorded rank crossings in a period far too short for editorial change. A drop inside the noise floor of the measurement is not evidence about your pages.

Which AI visibility numbers are repeatable?

The ones measured on your own server's response. Whether robots.txt allows a crawler token, what text a crawler receives compared with what a browser renders, and what structured data a page carries are all properties of bytes you control, so two scans agree unless you changed something. Numbers describing what an answer engine said are samples and should be reported with a date and an engine name.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.