BlogFindings

llms.txt vs sitemap.xml: a median of 23 URLs against 969 on the 239 sites publishing both

Lantad requested /llms.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026 as LantadBot, following redirects and executing no JavaScript. 305 answered HTTP 200 with a text body that was not HTML, and a control path that cannot exist returned text on none of them. 298 were real files. On the 239 of those whose sitemap resolved completely, the llms.txt named a median of 23 URLs and the sitemap declared a median of 969.

17 min read Lantad

So this run counted. Lantad requested /llms.txt from all 1,419 hostnames in the two corpus seed files committed to this repository on 29 September 2026, as LantadBot, following redirects and executing no JavaScript, then went and fetched the robots.txt and the sitemap of every site that answered with a real one. The result is not that one file wins. It is that on almost every site publishing both, the llms.txt is a short list drawn out of an inventory the site was already publishing.

In short

  • Lantad requested /llms.txt from all 1,419 hostnames in this repository's committed corpus on 29 September 2026: 305 answered HTTP 200 with a non-HTML text body, and a control path that cannot exist returned text at HTTP 200 on 0 of those 305, so none of them is a soft 404.
  • On llms.txt vs sitemap.xml the question is size before it is anything else: across the 239 sites whose sitemap resolved completely on 29 September 2026, the llms.txt named a median of 23 URLs against the sitemap's median of 969, and the sitemap was the larger file on 227 of the 239.
  • The two files are mostly not rival inventories: on the 90 sites whose sitemap is a single file rather than an index, a median of 86.4 percent of the URLs named in the llms.txt already appeared in that sitemap on 29 September 2026.
  • The exception is the markdown convention, and on this corpus it is rare: 9 of those 90 sites had an llms.txt where at least half the targets end in .md, their median overlap with the sitemap was 0 percent, and 82 of the 298 publishers named a .md target at all.
  • 279 of the 298 llms.txt files began with an H1, which the llms.txt proposal modified on 10 August 2026 calls the only required section, and 5 of the 305 text responses were zero bytes long.
StepHostnamesNote
Hostnames requested1,419392 platform frame, 1,027 industry frame
Answered HTTP 200 with a non-HTML text body305257 text/plain, 47 text/markdown
Same answer for a control path that cannot exist0so none of the 305 is a soft 404
Returned a file of zero bytes5excluded from the findings below
Confirmed llms.txt publishers298non-empty, with an H1 or at least one link
Served a parseable robots.txt294of the 298
Declared a sitemap in that robots.txt277of the 298
Sitemap resolved completely, uncapped239the comparison set used below
What Lantad found when it requested /llms.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026 as LantadBot/1.0, redirects followed and no JavaScript executed, then fetched the robots.txt and first declared sitemap of every confirmed publisher.

llms.txt vs sitemap.xml: what is each file specified to do?

The two files were written for different readers and neither specification hides it. Google's sitemap documentation, which carried a last updated date of 8 July 2026 when it was read for this post, describes a sitemap as a way to tell search engines which URLs you prefer to show in search results, and states that all formats limit a single sitemap to 50MB uncompressed or 50,000 URLs. It is also blunt about what submitting one buys you: a sitemap is merely a hint, and it does not guarantee that Google will download the sitemap or use it for crawling URLs on the site.

The llms.txt proposal is now at v2, published on 3 September 2024 and modified on 10 August 2026. It asks for an H1 with the name of the project or site, and says in those words that this is the only required section. After that it wants a blockquote summary and zero or more H2 sections holding file lists, where each entry is a markdown hyperlink followed optionally by a colon and a note.

The proposal also argues directly against the file this post compares it to, and that argument is the thing worth testing. It says sitemap.xml is a list of all the indexable human-readable information available on a site, and that this is not a substitute for llms.txt because it will not list the LLM-readable versions of pages, will not include external URLs, and will generally cover documents that in aggregate are too large to fit in a context window.

Those are three falsifiable claims about real sites, not matters of taste. An AI crawler arriving at a domain can request both files in two round trips, and the field that argues about which to write calls itself generative engine optimization or answer engine optimization, so the question of which one to write is decided by what each actually contains once a site has finished writing it.

Criterionsitemap.xmlllms.txt
First published2005 as a protocol2024 as a proposal
FormatXML, machine generatedMarkdown, usually hand written
Stated purposeWhich URLs you prefer in search resultsA curated overview for a language model
Size ceiling in the spec50,000 URLs or 50MB uncompressednone stated
Required elementA urlset with at least one locAn H1, the only required section
Declared in robots.txt277 of 298 publishers didno discovery line exists
Median URLs named, measured here96923
Honoured by search enginesA hint, not a guarantee, per GoogleNo engine documents reading it
The two files set against the criteria a site owner actually applies when deciding which to publish. The specification column states what each document says of itself: Google's sitemap documentation as read on 29 September 2026 carrying a last updated date of 8 July 2026, and the llms.txt proposal v2 as modified on 10 August 2026. The measured column is what Lantad found on this corpus on 29 September 2026.

How many of 1,419 sites published an llms.txt?

305 of the 1,419 hostnames answered a request for /llms.txt with HTTP 200 and a body that was not HTML. That number on its own would be worth very little, because a large share of the web answers 200 to almost anything, and a text file that is really a fallback page would inflate every figure below it.

So the run controlled for it. Every one of the 305 hosts was asked a second time for a path that cannot exist on any real site, /llms-parity-control-9f3a2b.txt. Not one of the 305 returned a non-HTML text body at HTTP 200 for that path. That is a cleaner result than expected and it means the 305 are answering for the specific filename rather than for anything ending in .txt, which is the failure this site has measured in other guises, most recently where 232 of 1,059 robots.txt files carried a defect and 33 sites answered with a web page. The same control is worth running by hand before trusting any file a site claims to serve, which is what our robots.txt tester does for the other one.

Five of the 305 returned a file of zero bytes: letterhunt.co, listr.pro, cnrs.fr, hel.fi and mediclinic.co.za. An empty file is a deliberate-looking answer that carries nothing, and all five are excluded from every figure after this paragraph. That leaves 298 confirmed publishers, where confirmed means the body was not empty and it either opened with an H1 or contained at least one link.

Conformance to the proposal is high but not total. 279 of the 298 began with an H1, so 19 did not, among them salesforce.com, twilio.com, usbank.com and nsw.gov.au. 26 of the 298 contained no link at all, which makes them a description of a site rather than an index into one; edx.org, expedia.com, metlife.com and natwest.com are in that group. The rate itself is worth stating plainly against the 23 llms.txt files this site found across 200 hostnames on 7 September 2026: a bigger and differently drawn corpus, a broadly similar share.

  • Control returned text at 200 0 of 305 No host in this set answers text to any path, so the 305 are real responses to the filename.
  • Zero byte file 5 of 305 letterhunt.co, listr.pro, cnrs.fr, hel.fi, mediclinic.co.za. Excluded from every later figure.
  • Confirmed publishers 298 Non-empty, and carrying either an opening H1 or at least one markdown link.
  • Opened with an H1 279 of 298 The proposal calls the H1 the only required section, so 19 files omit the one thing required.
  • Contained no link 26 of 298 A summary with no file list. It describes the site without indexing anything on it.
  • Named at least one .md target 82 of 298 The markdown convention the proposal rests its case on is present on just over a quarter.
How the 305 HTTP 200 text responses to /llms.txt were classified by Lantad on 29 September 2026. The control is a request for /llms-parity-control-9f3a2b.txt sent to the same host, which no real site can serve, so a text answer there would mean the llms.txt result was a fallback rather than a file.

A median of 23 URLs against a sitemap's 969

For every one of the 298 publishers the run then fetched /robots.txt, took the first Sitemap line it declared, and fetched that. 294 returned a parseable robots.txt and 277 declared at least one sitemap in it. 285 sitemaps came back at HTTP 200 as XML, 273 of them from a declared line and 12 found by asking for /sitemap.xml on a site whose robots.txt named none.

180 of those 285 were sitemap index files rather than URL lists, which is the shape a large site is supposed to use once it passes Google's 50,000 URL ceiling. Each index was resolved by fetching its children, up to a cap of 25 children per site. 32 sites hit that cap and are excluded rather than reported as a floor, which leaves 239 sites where the sitemap total is a complete count rather than a partial one. Every figure in this section is drawn from those 239.

The llms.txt named a median of 23 URLs. The sitemap declared a median of 969. The sitemap was the larger of the two on 227 of the 239 sites, and the median per-site ratio was 32.7 to one. That ratio is in the range this site found when it counted a sitemap against a link crawl of the same sites, where six sites exposed 985 linked paths and declared 200,712 URLs between them. The quartiles say the same thing more precisely: a quarter of the llms.txt files named 3 URLs or fewer while a quarter of the sitemaps declared 198 or more, and at the top the gap becomes absurd. n8n.io published an llms.txt naming 4 URLs beside a sitemap declaring 336,193. canadiantire.ca named 12,227 against 289,259.

12 sites inverted it, and they are worth naming because they are the shape that argues for the file. usbank.com named 3,843 URLs in its llms.txt against 3,336 in its sitemap, and salesforce.com named 3,677 against 1,242. On those sites the hand-written file is not a summary of the sitemap. It is a larger claim about what exists than the sitemap makes, which is a different problem and not obviously a better one, since the sitemap is the document a search engine actually reads.

  • Sitemap, upper quartile 3287 URLs
  • Sitemap, median 969 URLs the same sites' llms.txt median is 23
  • Sitemap, lower quartile 198 URLs
  • llms.txt, upper quartile 74 URLs
  • llms.txt, median 23 URLs
  • llms.txt, lower quartile 3 URLs
URLs named by each file across the 239 sites measured by Lantad on 29 September 2026 whose sitemap resolved completely. Sites whose sitemap index exceeded the 25 child cap are excluded rather than reported as a floor, so every count here is complete.

Are the URLs in an llms.txt already in the sitemap?

Size is the easy half. A short file is not redundant if it points somewhere the long file does not, so the run compared the two sets of URLs directly rather than only their counts.

That comparison needs a complete sitemap URL set, and resolving an index of unknown depth makes completeness a judgement call. So it was restricted to the 102 sites whose first sitemap was a single urlset rather than an index, where one request returns the whole set and nothing is capped or inferred. 90 of those compared cleanly; the remaining 12 either had no usable loc elements or no link in the llms.txt. URLs were compared after normalising away the scheme, a www prefix, a query string, a fragment and a trailing slash, so a cosmetic difference does not read as a different page.

On the median site, 86.4 percent of the URLs named in the llms.txt were already in the sitemap. 14 of the 90 sites had complete overlap, where the llms.txt did not name a single URL the sitemap had missed. Whether those URLs answer is a separate matter, and this site has counted that too: 31 llms.txt files held 5,567 links of which 75 were dead. Pooled across all 90 sites the figure is lower, 12,385 of 20,778 URLs or 59.6 percent, and the gap between the pooled and the median figure is itself the finding: a small number of very large llms.txt files drag the pooled number down while the typical site's file is almost entirely contained in its sitemap.

So for most publishers the answer to the comparison is not that one file beats the other. It is that the llms.txt is a shortlist, drawn from an inventory the site already publishes in full, and that the editorial act of choosing 23 URLs out of 969 is the only thing the newer file adds. That may well be worth doing. It is a different claim from the one the proposal makes, and it is worth knowing which of the two you are making before you spend an afternoon on the file. This site has published what the evidence says about whether llms.txt is fetched at all, and that finding has not moved.

MeasureValueWhat it means
Sites compared90Single urlset sitemaps only, out of 102 candidates
Median overlap per site86.4%The typical llms.txt is a subset of the sitemap
Sites with complete overlap14Every URL in the llms.txt was already declared
Sites with zero overlap11The two files name entirely different resources
Pooled across all URLs59.6%12,385 of 20,778, pulled down by a few large files
Median URLs, llms.txt40On this 90 site subset
Median URLs, sitemap792On the same 90 sites
Overlap between the URLs named in an llms.txt and the URLs declared in the same site's sitemap, measured by Lantad on 29 September 2026 across the 90 sites whose sitemap is a single file rather than an index, so the sitemap set is complete in one request. URLs were compared after normalising scheme, www, query, fragment and trailing slash.

The nine sites where the two files describe different things

The proposal's strongest argument is that a sitemap cannot list the LLM-readable version of a page, because those markdown twins are not indexable human-readable documents and have no business in a sitemap. That argument is correct, and the data shows exactly where it applies and how narrow that is.

Split the 90 compared sites by whether at least half the targets in the llms.txt end in .md. 9 sites clear that line. Their median overlap with their own sitemap is 0 percent, and 7 of the 9 have literally no URL in common with it. crawlee.dev is the clearest case: 355 unique targets, every one of them a .md file, against a sitemap of 3,890 ordinary pages. plaid.com names 1,410 targets of which 1,335 are .md. linear.app names 142 of which 140 are. On those sites the two files are not competing, they are describing two different renderings of the same site, and publishing both is the only way to declare both.

The other 81 sites are the ordinary case, and their median overlap is 87.5 percent. Among the 78 sites naming no .md target at all it is 88.3 percent. Across all 298 publishers, 82 named at least one .md target and only 7 had an llms.txt where every target was one, so the convention the proposal rests its case on is present on just over a quarter of these files and dominant on very few.

That markdown twin is a real mechanism and this site has measured how well it is served: 34 of 718 home pages answered a request for markdown, and six of those sent HTML back anyway, and separately no AI crawler in that test used content negotiation to ask for it. The convention has a second file attached to it as well, and 32 of 44 llms.txt publishers in a documentation corpus also served an llms-full.txt. A site that publishes markdown versions has a reason to publish an llms.txt that a site without them does not have, because for the second site the file can only ever repeat what its sitemap already says.

At least half the targets are .md: 9 sites

  • Median overlap with the sitemap: 0 percent
  • 7 of the 9 share no URL with it at all
  • crawlee.dev: 355 targets, all .md, sitemap of 3,890
  • plaid.com: 1,410 targets, 1,335 of them .md
  • linear.app: 142 targets, 140 of them .md
  • Here the two files genuinely describe different resources

Under half the targets are .md: 81 sites

  • Median overlap with the sitemap: 87.5 percent
  • 14 of them have complete overlap
  • 78 sites name no .md target anywhere
  • Their median overlap is 88.3 percent
  • clickhouse.com: 448 targets, none of them .md
  • Here the llms.txt is a shortlist of the sitemap
The 90 sites compared by Lantad on 29 September 2026, split by whether at least half the link targets in their llms.txt end in .md. The overlap figures are medians of the per-site percentage of llms.txt URLs that also appear in the same site's sitemap.

What to check on your own site, and what this does not show

The practical order follows from the numbers rather than from a preference. Check that your sitemap resolves and is declared, because it is the file with an installed base and a stated size ceiling, and because 21 of the 298 publishers here had an llms.txt while declaring no sitemap in robots.txt at all. Then ask whether you publish markdown twins. If you do, an llms.txt is the only place to declare them and it earns its keep. If you do not, work out what your 23 URLs say that your 969 do not, and if the honest answer is nothing, you have written a shorter copy rather than a second document.

None of that tells you the file will be read. This site has never measured an engine fetching an llms.txt, and the published evidence on that question is not encouraging. What this run measured is narrower and firmer: what the files contain, on real sites, on one day.

The limits are worth stating in full. This is a single request per file per host on 29 September 2026, so a file that changed the next day is not tracked, and a host behind a challenge that let LantadBot through might refuse another agent. The corpus is a stratified sample of 1,419 hostnames drawn on 30 July and 3 August 2026 and it is not a random sample of the web, so the 298 publisher count describes this frame and not the internet. The overlap figures come from the 90 sites with a single file sitemap, which skews toward smaller sites, and the size comparison excludes the 32 sites whose index exceeded the child cap, which skews away from the largest. Both exclusions were made to keep the counts exact rather than approximate, and both push in directions worth remembering when reading the medians.

One thing this measurement cannot see at all is ranking or citation. It counts what two files declare. Whether declaring more of your site in a shorter file changes what an answer engine says about you is a different question, and our methodology is explicit about which of the two we grade. What is measured here is the input, and the input is what our research keeps returning to: an AI visibility problem starts with prose parity and with what a machine can read, and on 227 of 239 sites the longer of the two readable files was the one nobody wrote by hand.

The order in which to decide between the two files, drawn from what this measurement found rather than from a preference. The branch at the markdown question is where the 9 sites in the previous section separated from the other 81.

Written by

Lantad

Published .

Two files on a site claim to tell a machine what is worth reading. One of them has been a published protocol since 2005 and is understood by everything that crawls. The other is a 2024 proposal that a site writes by hand, and this site ships a tool that checks it. The obvious question is whether the newer file says anything the older one does not, and that turns out to be a question about counting rather than about opinion.

Common questions

llms.txt vs sitemap.xml: which one should I publish?

Publish the sitemap first. It has an installed base, a stated size ceiling of 50,000 URLs, and a discovery line in robots.txt that 277 of the 298 llms.txt publishers measured here were already using. Add an llms.txt when you publish markdown versions of your pages, because a sitemap cannot list those. On the 90 sites compared on 29 September 2026 a median of 86.4 percent of the URLs in the llms.txt were already in the sitemap.

Is an llms.txt just a smaller sitemap?

On most sites measured here, close to it. Across the 239 sites whose sitemap resolved completely on 29 September 2026, the llms.txt named a median of 23 URLs against the sitemap's 969, and the sitemap was larger on 227 of the 239. The exception is the 9 of 90 compared sites whose llms.txt points mostly at .md files, where the median overlap with the sitemap was 0 percent because those markdown documents do not belong in a sitemap.

Do search engines and AI crawlers read llms.txt?

No engine documents reading it, and this site has published the evidence that valid llms.txt files largely go unfetched. Google's own sitemap documentation, last updated 8 July 2026, is careful to say that even a sitemap is merely a hint and does not guarantee that Google will download it or use it for crawling. Neither file is a delivery mechanism, and this measurement counted what the files contain rather than what any crawler did with them.

How was this measured?

Lantad requested /llms.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026 as LantadBot/1.0, following redirects and executing no JavaScript, then requested a control path that cannot exist to rule out soft 404 responses. For every confirmed publisher it fetched robots.txt, took the first declared Sitemap line, and resolved that sitemap, capping index files at 25 children and excluding any site that hit the cap.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.