# 16 percent of the sources four AI search engines cited were AI-generated

> An audit posted to arXiv on 22 May 2026 put 712 queries to ChatGPT, Copilot, Gemini and Perplexity through their own interfaces, collected 26,266 cited URLs and scraped 19,154 of them. A detection classifier labelled 3,056 of those, approximately 16 percent, as AI-generated, ranging from 7.3 percent of ChatGPT's cited sources to 27.8 percent of Copilot's. The authors present the figure as a lower bound, because the same classifier missed 31.4 percent of a sample of articles already documented as AI-generated.

- Canonical page: https://lantad.co/blog/sixteen-percent-of-cited-sources-were-ai-generated
- This file: https://lantad.co/blog/sixteen-percent-of-cited-sources-were-ai-generated.md
- Last substantive update: 2026-08-18

## Key facts

- **Published:** 2026-08-18
- **Category:** Findings
- **Author:** Lantad
- **Length:** 3202 words
- **Takeaway 1:** An audit of four generative search engines by Mowafak Allaham and Nicholas Diakopoulos, posted to arXiv on 22 May 2026 as arXiv:2605.23684, submitted 712 queries on politics, health and the environment through each engine's user interface and collected 2,848 responses.
- **Takeaway 2:** Of the 19,154 cited sources the authors successfully scraped, the Pangram classifier labelled 3,056, approximately 16 percent, as AI-generated: 2,916 as Highly Likely AI and 140 as Likely AI.
- **Takeaway 3:** The share differed by engine, at 27.8 percent of Copilot's cited sources, 14.7 percent of Gemini's, 9.4 percent of Perplexity's and 7.3 percent of ChatGPT's.
- **Takeaway 4:** The paper reports that Pangram classified only 55.2 percent of 105 articles from domains journalists had documented as AI-generated as Highly Likely AI, a false negative rate of 31.4 percent, which is why the authors call 16 percent a lower bound.
- **Takeaway 5:** Lantad ran no part of this audit and measures nothing about which sources an engine chooses to cite. Every figure below is read from the paper itself, fetched on 18 August 2026.

## Summary

Almost every argument for making a site machine readable ends at a citation. The page has to be reachable, the prose has to survive the fetch, and then an answer engine has to pick it. A study posted to arXiv on 22 May 2026 asks the question that sits one step past all of that: when a generative search engine does cite something, what is the thing it cited?

The paper is [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), by Mowafak Allaham and Nicholas Diakopoulos, published under a CC BY 4.0 licence. It put 712 queries on politics, health and the environment to ChatGPT, Copilot, Gemini and Perplexity, collected 26,266 unique cited URLs across 7,675 domains, scraped the text of 19,154 of them, and ran every one through an AI text detection classifier. Approximately 16 percent came back AI-generated.

Lantad measured no part of this and has never measured what any engine cites. What this scanner measures is whether [an AI crawler](https://lantad.co/glossary/ai-crawler) can reach a page and read the prose on it, which is the [method set out here](https://lantad.co/methodology) and a strictly earlier stage of the same pipeline. The audit is about the far end of that pipe, and it is worth reading precisely because it constrains what the near end can promise.

## What the audit submitted to each AI search engine, and what came back

The method is worth reading before any of the numbers, because two choices in it decide how far the findings travel.

The first is that the queries were submitted through each engine's user interface rather than its API. [The paper's methods section](https://arxiv.org/html/2605.23684v1) states that the authors built a Python scraper using Playwright to simulate a user, on the grounds that prior work has found responses to API requests differing from responses in the web interface, and that what a user actually sees is the thing worth auditing. That is the right call for the question being asked and it also means the sample captures the product rather than the model.

The second is where the queries came from. The politics and health queries were filtered out of the Search Arena dataset, which holds 24,069 real queries collected between 18 March and 8 May 2025 from 11,650 users across 136 countries. The authors excluded creative and malformed intents, reduced 1,325 candidates to those relevant to a United States context, and kept 432: 257 on politics and 175 on health. No environment queries survived that filter, so those came from the Climate Q&A dataset, 3,425 single turn queries collected in 2023 from users mostly in France, from which 280 were retained. The result is 712 queries that real people wrote, which is a genuine strength, and which is not the same as 712 queries written recently.

One date is missing and it matters enough to name. The paper gives the collection dates of both source datasets and does not state when the 712 queries were put to the four engines. The responses were therefore gathered somewhere between those datasets and the 22 May 2026 posting, and a reader cannot narrow it further from the paper. Anyone treating a figure here as a description of an engine's behaviour this month is adding a precision the authors did not claim.

What came back was 2,848 responses. Those responses cited 26,266 unique URLs across 7,675 unique domains, and a scraper built on newspaper3k successfully extracted body text from 19,154 of them, 72.9 percent, covering 6,258 domains or 81.5 percent of all domains seen. Everything downstream is computed on that 19,154, not on the full 26,266, which is a distinction the summaries of this paper tend to lose. Sample sizes of this kind are the reason [an AI visibility measurement needs a stated number of prompts](https://lantad.co/blog/how-many-prompts-an-ai-visibility-measurement-needs) before it means anything, and a separate study found that [visibility took seven runs per prompt to settle](https://lantad.co/blog/ai-visibility-took-seven-runs-per-prompt-to-settle). This audit ran each query once, which is appropriate for characterising a citation pool and would not be appropriate for ranking any individual brand's [generative engine optimization](https://lantad.co/glossary/geo) outcome.

## One in nine health answers carried no citation at all

Before the detection results, the paper reports something plainer that is easy to skip and harder to design around. Of the 2,848 responses collected, 2,596 carried in-text citations and 252 did not. That is 8.8 percent of answers with nothing to click.

The split by topic is not even. Health produced 120 of those uncited responses, the environment 92, and politics 40, which the paper gives as 11.6 percent of all health queries, 8.2 percent of environment queries and 5.7 percent of politics queries. The authors read the health figure, roughly one in nine, as evidence that these engines sometimes answer from what the model already holds rather than retrieving anything, and they raise the obvious consequence, which is that such an answer carries no timeliness guarantee and no source a reader can check.

For a site owner this is the floor under every readability improvement, and it is a floor nothing on the site can raise. A page that is perfectly fetchable, perfectly parsed and perfectly relevant is still absent from an answer that cited nobody. That is not a failure of the page. It is a property of the product, and it is one reason this site treats [AI visibility](https://lantad.co/glossary/ai-visibility) as an outcome it can influence rather than one it can promise. The platform notes for [getting cited in ChatGPT](https://lantad.co/how-to-get-cited/chatgpt) start from retrieval happening at all, and earlier work covered here found that [ChatGPT citations are rare and tend to land on the homepage](https://lantad.co/blog/chatgpt-citations-are-rare-and-land-on-the-homepage) when they happen.

There is a second reason to hold this number in view while reading the rest. Every percentage in the detection results is computed over sources that were cited, so the denominator excludes both the answers that cited nothing and every page that was never selected. That is the same denominator problem as the finding that [more citations did not mean more influence on the answer](https://lantad.co/blog/citation-count-is-not-answer-influence): what gets counted shapes what the count can mean.

## The 16 percent rests on one detector that missed a third of a known sample

The headline figure is a classifier output, so the classifier is part of the finding and the paper treats it that way at length.

Pangram, a transformer based classifier, sorted the 19,154 scraped sources into four buckets: Unlikely AI at 15,815 sources or 82.5 percent, Highly Likely AI at 2,916 or 15.2 percent, Possibly AI at 283 or 1.4 percent, and Likely AI at 140 or 0.7 percent. The authors then made a deliberately conservative choice and dropped the Possibly AI bucket entirely, on the grounds that it cannot be read as either co-authored content or as detector uncertainty, and that keeping it would inflate the estimate. Combining only Highly Likely AI and Likely AI gives 3,056 sources, the approximately 16 percent that the title of this post reports.

Then they tested the instrument twice. On a curated benchmark of 200 human texts drawn from the RAID collection and 200 texts the authors generated, both Pangram and GPTZero scored 100 percent, with no false positives. On a harder test they assembled 105 articles, fifteen each from seven domains that journalists had already documented as producing AI generated news, and Pangram classified 55.2 percent of them as Highly Likely AI, 16.2 percent as Mixed and 28.5 percent as Unlikely AI. Treating all 105 as genuinely AI-generated, which the authors call a reasonable assumption they cannot further validate, puts Pangram's false negative rate at 31.4 percent, against 30.4 percent for GPTZero. The two tools agreed on 82 percent of the sample and disagreed on 19 of the 105 articles.

The authors picked Pangram knowing it had the higher false negative rate of the two, precisely because a detector that misses more produces a lower estimate, and they wanted the lower bound. That is the opposite of the usual incentive, and it is the detail that makes the 16 percent worth quoting rather than discounting. It also sets a hard ceiling on how the figure may be used: it is one detector's floor on one corpus, not a census of the web, and the paper's own limitations section says another tool would yield a different proportion.

This is the same discipline that keeps a scan from issuing a confident grade on thin evidence, which is the argument in [why we withhold a grade](https://lantad.co/blog/why-we-withhold-a-grade), and the same one behind reporting that [hidden text was the weakest attack on AI search](https://lantad.co/blog/hidden-text-was-the-weakest-attack-on-ai-search) rather than the strongest. A number that names its own error rate is more useful than a number that does not.

## The AI-generated sources sit in the long tail, not in the head

The second half of the paper is about distribution, and it changes what the 16 percent means for anyone trying to be cited.

Citations are unequally spread but not narrowly so. Across all four engines the Gini index over cited domains is 0.68, and the 25 most cited domains account for 23.8 percent of all citations. The remaining 76.2 percent is spread very thin: 59.1 percent of domains were cited exactly once and a further 16.5 percent exactly twice. Per engine, ChatGPT concentrated hardest at a Gini of 0.648, then Gemini at 0.595, Perplexity at 0.563 and Copilot at 0.492, which is the same ordering as the hero table reversed and is worth noticing: the engine that spread its citations widest also cited the most AI-generated content.

The AI-generated sources are not in the head of that distribution. The paper reports that the top 25 domains account for only 2.9 percent of the AI-generated sources, with 97.1 percent falling among the rest, and separately that the top 35 domains ranked by AI-generated content hold 24.7 percent of it. At the domain level, 1,754 of the 6,258 analysed domains, 28.0 percent, had at least one source classified AI-generated, and of those 66.2 percent had exactly one, 17.7 percent had two, and 16.0 percent had three or more. So this is mostly not a story about a handful of content farms. It is a thin layer spread across a quarter of the domains an engine reaches for.

Two things follow for a site owner, and they point in opposite directions. The encouraging one is that the citation pool is genuinely wide, so being cited once is the ordinary outcome rather than a near miss, which fits the earlier finding that [a brand's own domain drew 2.9 percent of AI citations](https://lantad.co/blog/ai-citations-mostly-point-at-other-companies) while most pointed elsewhere. The discouraging one is that a wide pool selected with little apparent quality filtering is a pool your page competes in on terms you do not set. A controlled experiment covered here found [four factors decided the first citation and formatting was not one of them](https://lantad.co/blog/four-factors-decided-the-first-citation), and a separate evaluation found that [naming trusted domains moved citations from 12 to 21 percent](https://lantad.co/blog/trusted-domain-list-moved-citations-12-to-21-percent), which is a system operator's lever rather than a publisher's. The platform notes for [getting cited in Perplexity](https://lantad.co/how-to-get-cited/perplexity) describe the retrieval behaviour this distribution is a consequence of.

## What a site owner can do about AI search citations, which is less than it looks

The inconvenient reading is the one to lead with, because this paper offers a publisher almost no lever at all.

Nothing in it says that being human written improves your odds of being cited. It reports the opposite pattern: content a detector flags as machine generated was cited at a material rate by all four engines, and most heavily by the engine that spread its citations widest. There is no measurement here of whether AI-generated content was preferred, penalised or simply not assessed, and the authors do not claim one. A site that writes carefully is competing for the same slots as a site that does not, and this audit gives no evidence that the engines are separating the two. Writing well remains a good idea for readers. It is not, on this evidence, a ranking mechanism.

What the paper does confirm is where the readability question sits, which is strictly upstream of everything above. The researchers could not scrape 7,112 of the 26,266 cited URLs. They attribute part of that to format: 14.2 percent of the URLs they could not scrape were PDFs, and they separately report 16.2 percent of the overall sources as other modalities such as video, two shares written against different denominators in the paper itself. The rest they attribute to content being inaccessible or removed, per the HTTP response code. Their extractor was newspaper3k, which pulls body text from a fetched response and runs no JavaScript, and the paper does not name rendering as a cause of any failure. It is a gap rather than a finding, and it is exactly the gap a nearby study did not have, since the deep research benchmark covered in [links worked and facts did not](https://lantad.co/blog/deep-research-links-worked-facts-did-not) used an extractor that did render.

That difference is the whole reason this scanner fetches a page twice rather than once. The distance between what an origin returns to a plain client and what appears after scripts run is [prose parity](https://lantad.co/glossary/prose-parity), and it is invisible to anyone auditing from the outside with a single non-rendering fetch. Crawler documentation does not close the gap either: a count published here on 14 August 2026 found the words JavaScript, render, scroll, viewport and lazy appearing zero times across the crawler pages of OpenAI, Anthropic and Perplexity, which is the finding behind [scroll-loaded content never loading for a crawler that does not interact](https://lantad.co/blog/lazy-loading-and-crawlers-that-do-not-scroll). If you want to know what an unauthenticated non-browser client currently receives from one of your own URLs, that is an external observation and [what GPTBot sees](https://lantad.co/tools/what-gptbot-sees) is the shape of it.

So the honest summary is narrow. This audit tells you the pool you are competing in is wide, thinly filtered, and contains a measurable share of machine written pages that four major engines were willing to cite. It tells you nothing about how to enter that pool, and it tells you nothing about your own site. Being readable is a precondition and not a mechanism, which is the same boundary that keeps [the crawlability study](https://lantad.co/research/crawlability-study) reporting what a population of sites returns to a fetch rather than what any of them earned in an answer. The paper's stated limits are worth carrying with the numbers: English only, United States relevant only, single turn queries, and a single detector whose output the authors ask readers to treat as a lower bound.

## Questions and answers

**What share of AI search citations pointed at AI-generated content?**

Approximately 16 percent. Of the 19,154 cited sources the authors of arXiv:2605.23684 successfully scraped, the Pangram classifier labelled 3,056 as AI-generated, combining 2,916 Highly Likely AI and 140 Likely AI. The paper was posted on 22 May 2026 and covers ChatGPT, Copilot, Gemini and Perplexity.

**Did the four engines differ from each other?**

Substantially. AI-generated sources made up 27.8 percent of Copilot's cited sources, 14.7 percent of Gemini's, 9.4 percent of Perplexity's and 7.3 percent of ChatGPT's. The engine with the widest spread of citations, measured by a Gini index of 0.492 for Copilot against 0.648 for ChatGPT, was also the engine that cited the most AI-generated content.

**How reliable is the 16 percent figure?**

The authors present it as a lower bound. On a benchmark of 105 articles from domains already documented as AI-generated, Pangram classified only 55.2 percent as Highly Likely AI, a false negative rate of 31.4 percent. They chose that detector over GPTZero knowing it missed more, and excluded the ambiguous Possibly AI category, both of which lower the estimate rather than raise it.

**Does this mean a site should change what it publishes?**

Nothing in the paper supports that. It measures what four engines cited, not why, and takes no position on whether AI-generated content was preferred or penalised. The finding a site owner can act on is upstream of it: whether a page can be fetched and read at all, which the audit does not measure and which its own scraper failed on for 7,112 of 26,266 cited URLs.

---

Lantad measures whether AI crawlers can actually read a page: it fetches as a non-rendering
crawler, renders as a browser, and reports the gap. Free scan, one URL, no signup.

Method and weights: https://lantad.co/methodology | All pages as markdown: https://lantad.co/md | Crawler policy: https://lantad.co/bot
