BlogFindings

A trusted domain list moved AI citations from 12 to 21 percent

A pre-launch expert evaluation of Evrópuvefur, a government-funded Icelandic service answering questions about the EU, compared a curated corpus against open web search across 449 answers. Naming roughly sixty trusted domains in the system prompt raised the share of citations landing on one of them from 12 percent to 21 percent, and the country's most widely used news source was never cited across 287 web-search answers.

15 min read Lantad

Curated retrieval versus open web search in public AI information services, arXiv 2607.05217, submitted on 6 July 2026 and revised on 7 July, reports a pre-launch expert evaluation of Evrópuvefur, an independent, government-funded service run by the University of Iceland that answers questions about the European Union. The evaluation was conducted as Iceland prepared for its referendum of 29 August 2026 on whether to resume EU accession talks. Everything below is read from that paper. Lantad ran no part of it, holds no citation corpus, and has never measured which sources any assistant selects. What this post can add, and what the closing section does, is set the finding against the layer this site does measure, which is whether an AI crawler can fetch and read a page at all, and what that means for anyone treating generative engine optimisation as a matter of persuading a model.

In short

  • A companion prompt ablation in arXiv 2607.05217, submitted 6 July 2026, reports that listing trusted domains in the system prompt raised the share of citations landing on a listed domain from 12 percent of 1,330 citations to 21 percent of 1,183.
  • Five reviewers produced 551 evaluations covering 449 distinct answers, split 262 retrieval-augmented answers against the curated corpus and 187 answers from open web search, and recorded 128 source flags between them.
  • At least one cited source was flagged in 35 percent of reviewed web-search answers, being 65 of 187, against 6 percent for the curated path, being 16 of 262; the 110 web flags were 87 for untrustworthiness and 23 for irrelevance, while all 18 curated flags were for being out of date.
  • The curated path cited better sources and answered far less: 48.5 percent of its answers addressed the question against 91.5 percent for web search, and it left between a third and three-fifths of questions unanswered depending on question type.
  • Lantad measured none of this and holds no citation corpus. The study's relevance here is that both retrieval paths sit downstream of whether a page can be fetched and read at all, which is the only layer a scan touches.
MeasureCurated corpus (RAG)Open web search
Answers reviewed262187
Answers with a flagged source16 (6%)65 (35%)
Source flags recorded18110
What the flags saidOut of date only87 untrustworthy, 23 irrelevant
Addressed the question48.5%91.5%
The two retrieval paths as evaluated in arXiv 2607.05217, submitted 6 July 2026. Figures reported from the paper, not measured by Lantad.

What the evaluation actually measured

The design is simple enough to describe in a sentence and that is what makes it useful. The same service was run down two different retrieval paths, a curated local corpus and open web search, and domain experts scored the answers that came out of each. Five reviewers produced 551 evaluations covering 449 distinct answers, being 262 from the retrieval-augmented path over the curated corpus and 187 from web search, and recorded 128 source flags across them. Reviewers scored each answer against a seven-criterion rubric and, separately, flagged individual cited sources, which is the split that makes the result readable: an answer can be fluent, on topic and well scoped while resting on a source that should not have been used.

The seven criteria were whether the answer answers the question, whether it is factually accurate, whether it draws on relevant sources, whether it is free of hallucinations, whether it keeps to appropriate scope, whether it reads well in Icelandic, and whether it is publishable with only minor edits. Source flags were a separate axis. For web sources a reviewer could flag untrustworthiness or irrelevance; for the curated corpus the only flag available in practice was that a source was out of date.

The finding that matters most for anyone selling or buying visibility work is stated in the abstract and is not a matter of interpretation: fluency and topical fit did not predict source trustworthiness. The qualities a reader can see in an answer carry no information about the quality of what the answer was built from. That is a measurement problem before it is a content problem, and it is the same distinction we have written about from the other side, where a citation count is not answer influence and counting appearances tells you less than it appears to.

It is worth being exact about what kind of system this is, because the temptation is to generalise it to ChatGPT and Perplexity and stop there. Evrópuvefur is a public information service with a defined subject, a known corpus and an operator who can change the prompt, which is precisely why the ablation below is possible at all. Nobody can run that experiment on a commercial assistant from the outside. The paper does not claim its numbers transfer, and neither does this post. What transfers is the structure of the problem: a retrieval step selects a small set of documents, and everything a reader eventually sees is downstream of that selection. Any account of AI visibility that skips the retrieval step is describing the last mile and calling it the journey. The full paper, including the rubric and the ablation, is at arXiv 2607.05217.

The two retrieval paths compared in arXiv 2607.05217, drawn from the paper's description of the evaluation. A diagram of the design described in the paper, not a measurement of any site.

Web search answered more questions and cited worse sources

The headline comparison is a trade rather than a winner. Open web search answered far more of what it was asked: 91.5 percent of its answers addressed the question, against 48.5 percent for the curated path, a gap the paper reports at p below 0.001. Set against that, at least one cited source was flagged in 35 percent of reviewed web-search answers, being 65 of 187. The curated path was flagged on 6 percent, being 16 of 262.

The composition of the flags is the part worth sitting with. Of the 110 flags recorded against web sources, 87 were for untrustworthiness and 23 for irrelevance. Every one of the 18 flags against curated sources was for being out of date. Those are not the same kind of defect. A stale source is a maintenance problem with a known fix and a known blast radius. An untrustworthy source is a judgement failure at the moment of selection, and the reader has no way to see it, because the answer reads the same either way.

The curated path bought its accuracy by declining. The paper reports that it left between a third and three-fifths of questions unanswered depending on question type, highest for discourse positions at 59 percent and historical context at 55 percent, and lowest for referendum-specific questions at 31 percent. Among the curated answers that failed the criterion of answering the question, 87 percent carried a disclaimer saying so. A system that says it does not know is behaving well, and it is also a system that a great deal of published guidance about being cited would score as broken.

Volume, meanwhile, was not the constraint on the web path. Across the reviewed web-search answers the model cited 1,088 sources covering 442 distinct URLs, a mean of 5.8 per answer. There was no shortage of citation slots. The problem was which pages filled them, which is the same thing a controlled study found when it isolated what decides a citation and reported that four factors decided the first citation, and formatting was not one. It also sits alongside the observation that a brand's own domain drew a small minority of AI citations in a separate corpus: the slots exist, and they mostly go somewhere else. For anyone working on a specific surface, our notes on getting cited by Perplexity and the reference material on answer engine optimisation describe what is and is not within a publisher's control here.

  • Web search: at least one flagged source 35% 65 of 187 reviewed answers
  • Curated corpus: at least one flagged source 6% 16 of 262 reviewed answers
  • Web search: addressed the question 91.5% p below 0.001 against the curated path
  • Curated corpus: addressed the question 48.5% declines when the corpus falls short
Answers with at least one flagged source, and the share of answers addressing the question, per arXiv 2607.05217. Reported from the paper, not measured by Lantad.

A trusted domain list in the prompt moved citations from 12 to 21 percent

The companion ablation is the result with the widest application, because it tests the thing almost everyone reaches for first. The operators put a list of trusted domains into the system prompt, roughly sixty of them across EU institutions, EEA and EFTA bodies, Icelandic government agencies, international organisations, foreign governments, reputable news outlets and scholarly sources, and then measured where citations actually landed. The paper's sentence is exact: with the list in the prompt, 21 percent of the 1,183 citations landed on a listed domain; with the list removed, 12 percent of 1,330 did.

Read that in the direction that matters. The operator of the system, with full write access to the system prompt and a curated list of exactly the domains it wanted, moved the outcome by nine percentage points and still saw seventy-nine percent of citations land somewhere else. The paper's own summary of this is that prompt-level steering is weak. If the party running the model cannot reliably direct its citations by asking it to, the notion that a publisher can do so from outside, by writing in a particular register or adding a particular file, needs more evidence than it usually gets offered.

This is the same shape as a result we covered earlier, where the most-discussed manipulation turned out to be the least effective and hidden text was the weakest attack on AI search. Instructions aimed at the model, whether written by an attacker in invisible markup or by an operator in a system prompt, keep underperforming relative to how much attention they attract. What moves retrieval is what is in the corpus and what the retriever can reach, and reaching is decided earlier, by whether a compliant client is permitted to request the page at all under the robots exclusion protocol and by what the origin returns when it does.

There is a second reading worth naming, which is about where citations concentrate rather than how many there are. Our own note that ChatGPT citations are rare and land on the homepage describes a related pattern: retrieval collapses onto a small number of well-established addresses, and the long tail of individually excellent pages does not get sampled evenly. A domain list in a prompt is an attempt to override that concentration by instruction. It shifted it, partially. It did not replace it, which is roughly what you would expect if the underlying driver is entity confidence and the retriever's prior about which hosts are worth reading, rather than the wording of the request.

Prompt conditionCitations countedLanded on a listed domain
Trusted domain list present1,18321%
Trusted domain list removed1,33012%
Differencen/a9 percentage points
The companion prompt ablation in arXiv 2607.05217: citations landing on a domain named in the system prompt. Reported from the paper, not measured by Lantad.

The most used news source in the country was never cited

One sentence in the paper does more damage to the standard model of AI visibility than the whole ablation. Across all 287 web-search answers, the system never cited RÚV, the public broadcaster and the country's most widely used news source. Not rarely. Not less than expected. Zero times, across every web-search answer the service produced.

It is worth being careful about what that does and does not establish, because it is the kind of fact that invites overreach. The paper reports the absence; it does not publish an audit of why. There is no measurement here of whether the broadcaster's robots.txt permitted the retriever, whether its pages returned content to a non-browser client, whether the search backend surfaced it and the model passed it over, or whether the corpus simply never contained it. Each of those is a different failure with a different owner, and the paper's finding is that the outcome was zero regardless of which one caused it.

That ambiguity is the point for a site owner rather than a reason to dismiss the finding. A publisher looking at their own absence from an assistant's answers faces exactly this diagnostic problem, and the available explanations are not equally actionable. Some are on the retrieval side and cannot be reached from your server at all. Some are entirely on your side and are ordinary configuration. We have measured the shape of the second category before: blockers of Google-Extended were retrieved less by AI Overviews, and the pattern of who blocks AI crawlers splits by credibility in ways that mean the outlets most worth citing are sometimes the ones least reachable.

The general lesson is that authority in the human sense does not transfer automatically to retrieval. Being the most read source in a country is a fact about people. Being retrievable is a fact about servers, tokens and status codes, and the two are only loosely coupled. That decoupling is the entire premise of measuring the machine-facing side separately, and it is why structured data and a clean markup story are necessary without being sufficient: they help a system that has already fetched your page, and they do nothing for one that never did. If your interest is a specific assistant, our page on getting cited by ChatGPT sets out which of these levers a publisher actually holds.

Sample Illustrative, not a measurement of any real site.

  • Robots.txt disallow Checkable Resolved per crawler token against your own file. Entirely on the publisher's side and cheap to verify.
  • Origin or CDN refusal Checkable A 403 or a challenge served to a non-browser client ends the fetch before any content question arises.
  • Text absent from raw HTML Checkable A page whose body arrives only after JavaScript runs gives a JS-blind client nothing to quote.
  • Not in the curated corpus Not reachable Where retrieval runs over a fixed local corpus, inclusion is an operator decision no publisher setting affects.
  • Retriever ranked it below the cut Not reachable The search backend surfaced other documents. Not observable from outside the system.
  • Model passed it over Not reachable Selection among retrieved candidates. The ablation suggests prompt-level steering moves this weakly.
Candidate explanations for a source being absent from an assistant's citations. The paper reports the absence of RÚV across 287 web-search answers and publishes no cause; the causes below are the diagnostic categories, not findings of the study.

What this changes for whether an AI crawler can read your site

Nothing about the scan, and that is the honest answer rather than a deflection. This study measures a stage that begins after retrieval has already succeeded. A scan measures whether retrieval can succeed. They are adjacent and they are not the same question, and conflating them is how a great deal of visibility advice ends up unfalsifiable.

The useful consequence is a rule about ordering. If a system retrieves from a curated corpus, nothing on your server changes whether you are in it, and no amount of on-page work substitutes for the operator's inclusion decision. If a system retrieves from the open web, then the fetch has to succeed before any of the selection effects in this paper can even apply to you, which means the permission and delivery layers run first. That is what makes them worth fixing first: they are the part that is both decisive and yours. Whether the text a person reads is present in the raw HTML response is what prose parity scores, and what GPTBot sees will show it for a single URL without requiring you to take anyone's word for the mechanism.

The second consequence is about evidence standards, which is really what the paper is arguing for. Its closing claim is that source trustworthiness is a measurable yet largely invisible dimension of information quality, and that invisibility is the problem: nobody was looking, so nobody knew. The equivalent in this field is that most claims about what makes a page citable are made without a measurement attached and without naming what was not measured. When we surveyed the literature we found the same gap, that most GEO studies stop at the crawling stage and infer the rest. Our scoring method states what is weighted and why, and the research page publishes what our own sample does and does not support, including where it is too small to say anything.

The practical reading for a site owner is short. Treat retrieval as a precondition you can test and citation as an outcome you cannot control, and spend accordingly. Verify that the crawler tokens you care about are permitted, that your origin answers them with a 200 rather than a challenge, and that the page text exists before JavaScript runs. Then accept that the selection step on top is governed by factors this paper shows are weakly responsive even to the operator's own instructions. A publisher who has done the first part has removed every reason for absence that was theirs to remove. A publisher who skips it and writes for the model is optimising a stage their pages may never reach, which is the more expensive mistake because it is invisible: it produces no error, no log line and no signal that anything is wrong.

  • Robots.txt resolution per crawler token Decided before any fetch, and measured by a scan. The paper's findings all occur after this point has been passed.
  • Status code from origin or CDN Decided at request time, and measured by a scan. A refused request produces no candidate document to select from.
  • Page text present in the raw HTML Decided by how the page is rendered, and measured by a scan. A client that does not execute JavaScript quotes what it received.
  • Presence in a curated corpus An operator decision. Not measured by any scan and not reachable from a publisher's own configuration.
  • Which retrieved source the model cites Not measured by Lantad. The paper's ablation moved this from 12 percent to 21 percent using the operator's own system prompt.
Where the study's findings sit relative to what a Lantad scan reads. Compiled from arXiv 2607.05217 and the scan layers described on the methodology page. Not a scored check, and not weighted.

Written by

Lantad

Published .

Most advice about being cited by an AI system assumes a single shape for how the citation happens: the system searches the open web, finds your page, reads it, and quotes it. That assumption is doing a lot of work, and it is worth checking against a case where somebody wrote down what the machine actually retrieved. A paper published on arXiv gives one, and it is unusually well documented because the service it evaluates is public infrastructure rather than a product with a marketing incentive.

Common questions

Does adding a trusted domain list to a prompt make an AI cite those domains?

Only weakly, on the one published measurement of it. The companion ablation in arXiv 2607.05217 reports that with a list of roughly sixty trusted domains in the system prompt, 21 percent of 1,183 citations landed on a listed domain, against 12 percent of 1,330 with the list removed. The operator held full control of the prompt and still saw about four in five citations land elsewhere.

Is curated retrieval better than open web search for AI answers?

The study reports a trade rather than a winner. The curated path was flagged for source problems on 6 percent of answers against 35 percent for web search, and its only flag category was being out of date. It also addressed the question in 48.5 percent of answers against 91.5 percent for web search, and left between a third and three-fifths of questions unanswered depending on question type.

Why was the country's main news broadcaster never cited?

The paper reports the absence and does not publish a cause. Across all 287 web-search answers the system never cited RÚV, the public broadcaster and the country's most widely used news source. Whether that came from robots.txt, an origin refusal, absence from the retrieval index or the model's own selection is not established by the study, and those are different failures with different owners.

Did Lantad measure any of this?

No. Every figure in this post is read from arXiv 2607.05217, submitted on 6 July 2026 and revised on 7 July. Lantad ran no part of the evaluation, holds no citation corpus, and has never measured which sources any assistant selects. What a Lantad scan measures is the earlier layer: whether a named crawler is permitted, whether the origin answers it, and whether the page text is present in the raw HTML.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.