# A brand's own domain drew 2.9 percent of AI citations

> A preprint posted to arXiv on 18 June 2026 by a single author at the AI visibility vendor Ranqo breaks 149,912 citations from five AI engines into nine source classes. The brand's own domain accounts for 4,392 of them, 2.9 percent, while 75.2 percent point at other companies in the same category.

- Canonical page: https://lantad.co/blog/ai-citations-mostly-point-at-other-companies
- This file: https://lantad.co/blog/ai-citations-mostly-point-at-other-companies.md
- Last substantive update: 2026-08-10

## Key facts

- **Published:** 2026-08-10
- **Category:** Findings
- **Author:** Lantad
- **Length:** 3672 words
- **Takeaway 1:** A preprint posted to arXiv on 18 June 2026 as arXiv:2606.20065, written by Pratyush Kumar of the AI visibility vendor Ranqo, reports 102 brands, 3,508 completed tracking runs, 102,025 prompt responses and 149,912 source citations across five AI engines, gathered between March and May 2026 from that vendor's own production database.
- **Takeaway 2:** Of those 149,912 citations, the paper's section 6.5 table records 4,392 pointing at the tracked brand's own domain, which is 2.9 percent, against 112,763 or 75.2 percent pointing at commercial pages owned by other businesses in the same category.
- **Takeaway 3:** First-run visibility on unbranded prompts ran 72.9 percent for the 11 largest brands in the cohort, 43.6 percent for 36 mid-market brands and 11.4 percent for 55 niche brands, with 95 percent confidence intervals of 60.1 to 84.2, 36.4 to 50.9 and 4.2 to 20.3 respectively.
- **Takeaway 4:** The same paper states that its brand tiers were hand-coded from Wikipedia coverage, press and funding, that those signals are themselves proxies for web prominence, and that the ladder should therefore be read as sizing an expected effect rather than proving that prominence drives citation.
- **Takeaway 5:** Lantad measured none of this. Lantad ran no prompts against any engine, holds no citation corpus, and has never tested whether crawler readability correlates with citation rate on any cohort of sites.

## Summary

A scanner that tells you whether an AI crawler can read your website is answering a narrower question than the one most people arrive with. The question they arrive with is whether AI will recommend them. Those two things are connected, but the connection is an assumption rather than a measurement, and it is worth reading the studies that measure the far end of it even when, especially when, they are inconvenient for the near end.

This post reports one such study. It is a preprint, it has one author, and that author works for a company that sells the thing the paper measures. All three of those facts belong in the first paragraph rather than a footnote. What the paper has that most writing about [AI visibility](https://lantad.co/glossary/ai-visibility) does not is a large body of recorded engine output, a published breakdown of where citations landed, and an unusually direct statement of what its own numbers cannot support. Lantad has not measured any of it, ran none of these prompts, and operates no engine tracking. Everything below is reported from that paper and attributed to it.

## What the study measured, and who ran it

The paper is titled Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines. It was [posted to arXiv on 18 June 2026](https://arxiv.org/abs/2606.20065) as arXiv:2606.20065, runs to 14 pages with 4 tables, carries a CC BY 4.0 licence, and describes itself in its own comments field as a v1.0 preprint. It has a single author, Pratyush Kumar, and the affiliation printed under his name is Ranqo.

Section 6 states the dataset in one sentence: 102 brands, 3,508 completed tracking runs, 102,025 prompt responses, roughly 15,815 brand mentions and 149,912 source citations across five engines, gathered between March and May 2026. It also states where the data came from, which is Ranqo's production database. Ranqo sells AI visibility tracking. This is therefore a vendor publishing figures drawn from its own customers about the market it sells into, and that is a conflict worth naming before quoting a single number from it.

Naming it is not the same as dismissing it. The paper concedes the point itself in its limitations section, where the second limitation is headed convenience-sampled cohort and states that the brands skew toward SaaS, retail-execution, fintech and Indian DTC, and that the authors do not claim category representativeness. A vendor dataset with a stated skew is more useful than an anonymous one with no stated skew, and the survey of the field this blog covered when [45 GEO studies were reviewed and none showed a stable cross-platform causal effect](https://lantad.co/blog/geo-survey-45-studies-crawling-stage) makes the case that the shortage in this literature is measurement of any kind rather than measurement without commercial interest.

The engine configuration matters more than it usually does. The paper measures five engines, each on one search-enabled production variant held fixed: ChatGPT on GPT-5, Gemini on Gemini-3, Perplexity on Sonar, Claude on Claude Sonnet and Grok on Grok-3. Their retrieval policies differ by design. Gemini is recorded as always grounded, Perplexity as always grounded and search-native, Grok as always-on retrieval, while ChatGPT and Claude carry a weekly cooldown per brand. The citation surface differs too, and this is the part to hold on to: citations are read from annotations.url_citation on ChatGPT, groundingChunks on Gemini, a search_results array on Perplexity, text-block citations on Claude, and a citations array plus inline markdown on Grok. Comparing citation counts across five engines rests on those five API fields meaning the same thing, which is an assumption the paper does not test. Its fourth limitation, headed platform opacity, says as much: every result is inferred from API outputs, and the authors have no direct view into how the engines ingest, retain or retrieve. That distinction between what an engine emits and what an engine did is the same one that separates [generative engine optimisation](https://lantad.co/glossary/geo) from [answer engine optimisation](https://lantad.co/glossary/aeo) as terms, and it is the reason both are harder to measure than ranking ever was.

## Why only 2.9 percent of the citations pointed at the brand's own domain

The section headed source composition is where the paper earns its title, and its own heading states the finding before any table does: corporate dominance is mostly third-party brands, not the brand's own pages. Across the 149,912 citations drawn from mention-bearing prompts, the largest class is corporate and third-party brand pages at 112,763 citations, or 75.2 percent. The class labelled brand-owned, meaning the tracked brand's own domain, holds 4,392 citations, or 2.9 percent. The paper states the comparison in one line: only 2.9 percent of citations point at the brand's own domain, while 75.2 percent point at other companies in the same space.

The definition of that dominant class is what makes the number readable, and it is easy to get wrong. The paper defines corporate and third-party brand pages as commercial websites owned by businesses other than the tracked brand, meaning competitors, peers and vendors in the same category, including their homepages, product and landing pages, and corporate blogs. So this is not a finding that engines prefer publishers or reference sites over commercial ones. It is a finding that when an engine answers a question about a brand, the pages it cites are overwhelmingly commercial pages belonging to that brand's competitors. The paper's explanation is that engines preferentially construct alternatives-style answers, and in those answers the dominant source is a set of peer-brand product pages.

One arithmetic point prevents a misreading that would otherwise be very easy to make. The paper's abstract says about 78 percent of citations go to corporate websites, which looks like it contradicts the 75.2 percent in the table. It does not. [The full text](https://arxiv.org/html/2606.20065v1) says explicitly that own and third-party corporate pages together are about 78 percent of all 149,912 citations, so the abstract's figure is the two corporate rows added, 75.2 plus 2.9. The headline number and the inconvenient one are the same sentence, and anyone quoting the 78 percent without the split has quoted the part that flatters the reader.

Beneath the corporate classes the ordering is itself a correction. Video leads the non-corporate sources at 4.2 percent, ahead of tech and business media at 3.8 percent, community forums at 3.3 percent and Wikipedia at 2.6 percent. The paper notes that this reorders a common assumption, since forums are frequently reported as the leading non-corporate citation channel, and that Wikipedia is cited less often than video, media or forums. The page-type breakdown runs on two denominators that are worth keeping apart: the ranked listicle is 35.7 percent of content-classified citations and 21.0 percent of all citations, with generic articles at 31.0 and 18.3 percent. Quoting the first column as though it were the second inflates the figure by two thirds. None of this settles what a citation is worth once it exists, which is a separate question this blog has covered in [why a citation count is not answer influence](https://lantad.co/blog/citation-count-is-not-answer-influence), and it sits alongside the finding that [the page earning a ChatGPT citation and the page receiving the visit are frequently not the same page](https://lantad.co/blog/chatgpt-citations-are-rare-and-land-on-the-homepage).

## The visibility ladder ran 72.9, 43.6 and 11.4 percent by brand tier

The paper calls its headline finding the brand-stature visibility ladder, and it is a first-run measurement: what share of relevant unbranded prompts mention the brand at all on the very first tracking run, before anyone has done anything. Tier 1, with 11 brands, recorded 72.9 percent with a 95 percent confidence interval of 60.1 to 84.2. Tier 2, with 36 brands, recorded 43.6 percent with an interval of 36.4 to 50.9. Tier 3, with 55 brands, recorded 11.4 percent with an interval of 4.2 to 20.3. The abstract rounds these to 73, 44 and 11 and describes the gap as about 30 percentage points per step.

How the tiers were assigned is the part that decides what the ladder means. The paper states the rubric plainly: Tier 1 requires a high-coverage Wikipedia article plus recent top-tier press plus public or Series-C-or-later status, Tier 2 requires Wikipedia plus Series B or later funding or top-three-in-category regional leadership, and Tier 3 is everything else with an established web presence. A Tier 4 exists for single-run and test records and is excluded from this section. Assignment was hand-coded, and the paper says a formal inter-rater check is left to a later version.

Then comes the sentence that should travel with the figure everywhere it goes. The paper's third limitation states that because those tier signals are themselves proxies for web prominence, the ladder should be read as quantifying the size of an expected effect, not as independent proof that prominence drives citation. That is close to circular by construction: brands were sorted by how prominent they are on the web, and the finding is that the more prominent ones appear more often in answers assembled from the web. What the number adds is not the direction, which was never in doubt, but the magnitude, and the magnitude is large. The paper does report a robustness check, a Tier-1 leave-one-out reclassification giving a range of 70.5 to 76.7 percent, which is reassurance about the small 11-brand cell rather than about the rubric.

A companion table separates two things that get conflated constantly. Asked about by name, every engine recognised nearly every brand: 94.2 percent on ChatGPT, 94.0 on Gemini, 98.5 on Perplexity, and 100.0 on both Claude and Grok. Asked an unbranded question, the same brands surfaced far less: 22.1 percent on ChatGPT, 18.7 on Gemini, 23.9 on Perplexity, 51.5 on Claude and 12.0 on Grok. The engines are not failing to know these companies. They are declining to volunteer them, which is a retrieval and selection problem rather than a knowledge one, and it is the reason a brand-name spot check is close to useless as a visibility measurement. What a site can actually assert about its own identity, as opposed to what an engine chooses to do with it, is the narrower thing [entity confidence](https://lantad.co/glossary/entity-confidence) scores, built from [the structural signals that tell an AI who you are](https://lantad.co/blog/five-signals-that-tell-ai-who-you-are). The gap between being retrievable and being retrieved also runs through [the finding that nearly a third of AI Overview citations are not on page one](https://lantad.co/blog/ai-overview-citations-are-not-page-one), and it is why the practical guidance on [getting cited by ChatGPT](https://lantad.co/how-to-get-cited/chatgpt) is written around what a specific engine reads rather than around a general audience.

## Sentiment flipped 6.7 times more often than mention did

The most directly useful finding for anyone building or buying a visibility dashboard is in the section on measurement reliability, and it is about noise floors rather than about brands. For cells of brand, prompt and platform carrying at least three mentions on unbranded prompts, the paper reports how sentiment behaved across repeated runs: always positive in 29.9 percent of cells, a positive and neutral mix in 21.8 percent, always neutral in 2.3 percent, a negative and neutral mix in 0.5 percent, always negative in 0.0 percent, and flipping between positive and negative in 45.5 percent. The equivalent flipping rate for whether the brand was mentioned at all is 6.8 percent. Dividing one by the other is where the 6.7 times figure in [the paper's abstract](https://arxiv.org/abs/2606.20065) comes from, and both inputs are printed in the body, so the derivation can be checked rather than taken.

Two consequences follow, and the paper draws the operational one itself: a sentiment-weighted score needs a much larger denominator before it stabilises, and the platform in question uses at least 10 prompts per platform per brand. A dashboard that reports a sentiment trend from a handful of runs is reporting its own sampling noise with a direction attached. The second consequence is the paper's own sharper observation, which is that not a single cell was consistently negative, 0.0 percent, so negativity when it appears is transient rather than settled.

The honest reading requires the paper's sixth limitation. Sentiment there is classified by a model on the spans that mention the brand, so the 45.5 percent mixes genuine output variance with classifier noise, and the paper says explicitly that it has not separated the two. That makes 6.7 times an upper bound on real sentiment volatility rather than a measurement of it. It is still the right order of magnitude for a warning, and it is consistent with what a different and larger study found about repeat runs, covered here when [374,052 citations showed repeated runs of the same query sharing as little as 0.29 of their cited domains](https://lantad.co/blog/how-many-prompts-an-ai-visibility-measurement-needs).

Mention itself is comparatively stable, and the paper is careful to show that too rather than only asserting it. It reports 77.5 percent of cells as strictly deterministic and 93.2 percent as sitting outside the 30 to 70 percent band where a signal is genuinely undecided. Per-engine trajectories over repeated runs are mild: a mean slope of -1.34 percent per run on ChatGPT across 101 brands with an interval of -2.45 to -0.38, and -0.76 percent on Perplexity across 98 brands with an interval of -1.66 to -0.03, are the only two whose intervals exclude zero. Gemini at -1.09, Claude at -0.52 and Grok at +0.03 all have intervals crossing zero, which means no measured trend rather than a flat one. Reporting an interval that crosses zero as a decline is the same error as [printing a grade for a measurement that failed](https://lantad.co/blog/why-we-withhold-a-grade), and the discipline of naming an extracted signal as extracted rather than as established is the one this site had to build deliberately into [how rival names reach a report](https://lantad.co/blog/rival-names-are-extracted-not-entered).

## What this study measures, and what a crawler readability scan measures

It is worth being exact about what this paper looked at, because the answer is narrower than the subject suggests. It looked at engine outputs. It recorded which brands were named and which URLs were cited, read from five API response fields. At no point did it examine any property of the tracked brands' own websites. There is no robots.txt in it, no rendering check, no [structured data](https://lantad.co/glossary/structured-data) audit, no test of whether a named crawler is allowed to fetch a page. Its own fourth limitation says every result is inferred from API outputs.

What this scanner measures is the other side of the same system, and it is deliberately the side you control. Whether a named [AI crawler](https://lantad.co/glossary/ai-crawler) is permitted to fetch your pages is decided by your rules file and your edge. Whether the words a human reads are present in the HTML a crawler receives, rather than assembled afterwards by a browser, is [prose parity](https://lantad.co/glossary/prose-parity), and you can see the difference on a single URL with [the crawler-view tool](https://lantad.co/tools/what-gptbot-sees). Those are supply-side conditions, they are checkable from outside, and [how each is scored](https://lantad.co/methodology) is published.

The relationship between the two is a necessary condition, not a sufficient one, and the honest version of that sentence is uncomfortable for a company selling scans. A page a crawler cannot read cannot be cited. A page a crawler reads perfectly may still never be cited, and this paper is 149,912 citations' worth of evidence that for a brand in its third tier it usually is not: 11.4 percent first-run visibility, and 2.9 percent of citations landing on the brand's own domain. No amount of markup moves a brand from Tier 3 to Tier 1, because the tiers were defined by Wikipedia coverage, press and funding. Saying so is the same discipline as publishing that [97 percent of valid llms.txt files were never fetched](https://lantad.co/blog/what-the-evidence-says-about-llms-txt) while shipping an llms.txt tool.

The symmetry runs the other way too, and it is the reason neither measurement replaces the other. This study cannot tell a site owner whether their pages are readable, because it never looked at a page. A tracker watching engine outputs sees the effect and not the cause; a scanner reading your HTML sees a cause and not the effect. Anyone selling either as the whole picture is selling half of one. What a scan is genuinely for is removing the failure modes that are yours: text that never arrives in the HTML, a rules file that blocks a fetcher you meant to allow, a stack whose output depends on JavaScript running. That is the ground [six real storefront captures](https://lantad.co/blog/what-a-crawler-meets-on-a-real-storefront) cover, and it is a smaller claim than the market usually makes.

Finally, the gap this post does not close, stated so it is visible. Lantad has not measured whether crawler readability correlates with citation rate on any cohort of sites. Doing that honestly needs a joint dataset: scans of a set of domains on one side, tracked prompt outputs naming those same domains on the other, over a window long enough for the noise floors above to average out. Lantad holds the first half and none of the second, and [the research this site does publish](https://lantad.co/research) is scan-side only. Until somebody runs that study, the link between being readable and being cited is a reasonable mechanism and an unmeasured one, and it should be described that way by everyone selling against it, including us.

## Questions and answers

**How much of AI citations point at a brand's own website?**

In this study, 2.9 percent. A preprint posted to arXiv on 18 June 2026 as arXiv:2606.20065 breaks 149,912 citations recorded across five AI engines between March and May 2026 into nine source classes, and the class for the tracked brand's own domain holds 4,392 citations, or 2.9 percent. The largest class, at 112,763 citations or 75.2 percent, is commercial pages owned by other businesses in the same category. This is one vendor's cohort of 102 brands and it is not a measurement Lantad has taken.

**Does brand size determine whether AI mentions you?**

It correlated strongly in this cohort, and the paper declines to call it proof. First-run visibility on unbranded prompts was 72.9 percent for 11 top-tier brands, 43.6 percent for 36 mid-market brands and 11.4 percent for 55 niche brands. The paper's own limitation states that because the tiers were hand-coded from Wikipedia coverage, press and funding, which are themselves proxies for web prominence, the ladder sizes an expected effect rather than proving prominence causes citation.

**Why is AI sentiment tracking less reliable than mention tracking?**

Because it flips far more often between identical runs. The paper reports sentiment flipping between positive and negative in 45.5 percent of cells against 6.8 percent for whether the brand was mentioned at all, a ratio of 6.7 times. It also states that sentiment there is classified by a model and that genuine output variance has not been separated from classifier noise, so 6.7 times is an upper bound. Its operational conclusion is that sentiment-weighted scores need a much larger sample before they stabilise.

**If most citations go elsewhere, is checking AI crawler access still worth doing?**

Yes, but as a necessary condition rather than a sufficient one. A page a crawler cannot fetch or cannot read cannot be cited by anything downstream, and that failure is entirely within a site owner's control. Whether a readable page then gets cited depends on factors this study measured and a scan does not, chiefly brand prominence. Lantad has not measured any correlation between crawler readability and citation rate, and no such joint dataset is published here.

---

Lantad measures whether AI crawlers can actually read a page: it fetches as a non-rendering
crawler, renders as a browser, and reports the gap. Free scan, one URL, no signup.

Method and weights: https://lantad.co/methodology | All pages as markdown: https://lantad.co/md | Crawler policy: https://lantad.co/bot
