BlogFindings

Sitemap validator: 44 of 871 declared sitemaps delivered no file, and none of the rest broke the 50,000 URL limit

Lantad requested /robots.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 1 October 2026, then fetched the first sitemap each file declared. 1,104 answered HTTP 200 and 871 of those declared a sitemap. 827 of the declared sitemaps delivered a body, 38 answered a non-200 status, one timed out, and five named a value that is not an absolute URL.

20 min read Lantad

So on 1 October 2026 Lantad requested /robots.txt once from each of the 1,419 hostnames in this repository's two committed corpus seed files, as LantadBot, following redirects and executing no JavaScript, then fetched the first sitemap each file declared and checked what came back against the sitemaps.org protocol and against Google's sitemap documentation. 1,104 of the robots.txt requests answered HTTP 200, and 871 of those files carried at least one Sitemap line. The rest of this post works through what the 871 delivered, in the order of how much a failure costs the site, and you can run the robots.txt half of the same check against a single host with the robots.txt tester. None of this is evidence that any particular AI crawler read or failed to read these files, and the last section is about that gap rather than around it.

In short

  • A sitemap validator run on 1 October 2026 against the 871 robots.txt files that declared a sitemap found 827 delivering an HTTP 200 body and 44 delivering nothing a crawler could parse.
  • 24 of the 38 declared sitemaps that answered a non-200 status returned HTTP 403, among them congress.gov, oecd.org, tripadvisor.com, temu.com and revolut.com, each of which names a sitemap in robots.txt and then refuses to serve it.
  • Re-requesting all 39 unreachable sitemaps with a desktop Chrome user agent and with a Googlebot user agent on 1 October 2026 returned the identical status on 37 of them, so on those 37 the refusal is not a decision about which crawler is asking.
  • None of the 827 sitemaps that delivered a body exceeded the 50,000 URL or 52,428,800 byte ceilings the sitemaps.org protocol sets. The largest single urlset held exactly 50,000 URLs, at regeringen.se, and the largest body was 23,544,186 bytes, at salesforce.com.
  • 15 of the 827 sitemaps omitted the sitemaps.org 0.9 namespace declaration, and seven listed at least one loc value that is not an absolute URL, 312 entries between them, measured on 1 October 2026.
What the declared sitemap didSitesShare of 871
Delivered an HTTP 200 body82794.9%
Answered a non-200 status384.4%
Named a value that is not an absolute URL50.6%
Did not answer before the 45 second timeout10.1%
Exceeded the 50,000 URL ceiling00.0%
Exceeded the 52,428,800 byte ceiling00.0%
Measured by Lantad on 1 October 2026. One HTTPS GET for /robots.txt per hostname across the 1,419 hosts in worker/seeds/corpus-seeds-industry.json and worker/seeds/corpus-seeds-platform.json, sent as LantadBot/1.0, redirects followed, no JavaScript executed, then one GET for the first Sitemap value in each file. Shares are of the 871 files that declared a sitemap.

What does a sitemap validator actually check?

Two documents decide, and a validator should say which one it is quoting, because they are not the same kind of document. The first is the sitemaps.org protocol at sitemaps.org/protocol.html, which is the format definition and is where the famous limits live: "each Sitemap file that you provide must have no more than 50,000 URLs and must be no larger than 50MB (52,428,800 bytes)", and for an index, "Sitemap index files may not list more than 50,000 Sitemaps and must be no larger than 50MB (52,428,800 bytes)". The second is Google's sitemap documentation, carrying a last updated date of July 8, 2026, which states the same ceiling as "All formats limit a single sitemap to 50MB (uncompressed) or 50,000 URLs" and adds the instruction a validator can actually test on a per entry basis: "Use fully-qualified, absolute URLs in your sitemaps. Google will attempt to crawl your URLs exactly as listed."

The protocol is stricter than most summaries of it. It does not merely ask for absolute URLs, it bounds their length: "This value must be less than 2,048 characters." It does not merely ask that a sitemap belong to the site, it scopes a sitemap by its own location, saying "all URLs in a Sitemap must be from a single host" and giving the example that a file at http://example.com/catalog/sitemap.xml "can include any URLs starting with http://example.com/catalog/ but can not include URLs starting with http://example.com/images/". It also accepts formats people forget are legal, an "RSS (Real Simple Syndication) 2.0 or Atom 0.3 or 1.0 feed" and "a simple text file that contains one URL per line", which matters here because a validator that rejects everything without a urlset element would have failed several files in this sample that are not broken at all.

The one thing neither document governs is whether the sitemap arrives. That is robots.txt and HTTP, and it is the half of the problem this sample found. The Sitemap line is not scoped to a user-agent group under RFC 9309, which is why a declared sitemap reads as a public offer rather than a per crawler grant, a point this blog measured when 839 of 1,016 robots.txt files declared one. An offer that returns 403 is still an offer in the file and still nothing on the wire. The checks below run in that order for that reason: delivery first, then format. Our parsing rules are set out in the methodology.

CheckWhat it rests onWhat failing it costs
The declared value is an absolute URLsitemaps.org: the location must be a full URLThere is nothing to fetch
The sitemap answers HTTP 200Neither document; this is HTTPNo inventory reaches any crawler
The body is a urlset, a sitemapindex, a feed or a URL listsitemaps.org accepts all fourNothing is extracted
Under 50,000 entriessitemaps.org and Google both set itEntries past the limit are undefined
Under 52,428,800 bytes uncompressedsitemaps.org states the byte figureEntries past the limit are undefined
Every loc is a fully qualified absolute URLGoogle: crawled exactly as listedThat entry resolves against nothing
Every loc is under 2,048 characterssitemaps.org states the limitThe entry is out of spec
Entries sit within the sitemap's own locationsitemaps.org: a single host, scoped by pathOut of scope unless verified elsewhere
The checks run against each declared sitemap on 1 October 2026, with the document each rests on. Quoted text is verbatim from the source named in the middle column, read at source on the same date.

44 of 871 declared sitemaps delivered no file

This is the largest failure in the sample and it is invisible from a browser, because nobody opens their own sitemap in a browser. 38 of the 871 declared sitemaps answered with a non-200 status, one did not answer inside 45 seconds, and five named something that is not a URL a crawler can fetch. That is 44 sites, 5.1 percent of those that declared a sitemap at all, where robots.txt points a crawler at a URL inventory that does not exist on the wire.

24 of the 38 answered HTTP 403, and the list is not a list of small or neglected sites. congress.gov, imf.org, oecd.org, amsterdam.nl, gov.br, science.org, nejm.org, drugs.com, jamanetwork.com, tripadvisor.com, lufthansa.com, aa.com, temu.com, shein.com, rewe.de, laredoute.fr, bhphotovideo.com, revolut.com and museodelprado.es all name a sitemap in robots.txt and then refuse the request for it. Six answered 404, among them nyp.org, nursingworld.org, thejakartapost.com and vietnamairlines.com, which is the same outcome by a more honest route: the file was moved or never existed and robots.txt was not updated. Three answered 503, including washingtonpost.com and tokopedia.com. santander.com answered 500. heb.com answered 401. who.int answered 406. cloudsmith.com answered 429. connectbase.com answered 202, which is a success code carrying no sitemap, and spectator.co.uk timed out.

The five non-absolute declarations are a different mistake and a cheaper one to fix. cnrs.fr, deutsches-museum.de, home.cern and kosko.dev each write "Sitemap: /sitemap.xml", a root relative path, and bma.org.uk writes a bare "edinburghmeetingrooms.bma.org.uk/sitemap.xml" with no scheme. All five sitemaps very probably exist at the obvious address. The protocol requires a full URL in that field, and a crawler following it as written has no base to resolve against, so the line is a comment with extra steps. This is the same shape of defect as the ones in the robots.txt validator run from 19 September: a line that reads correctly to a person and resolves to nothing for a parser.

A 403 on a sitemap is also not the same event as a 403 on the home page. The status code a site returns when it cannot serve a file changes what a compliant crawler concludes, and under RFC 9309 a 404 and a 503 are opposites for robots.txt specifically. For a sitemap there is no such rule in either document, so a refused sitemap has no defined fallback at all. The crawler simply has no inventory and proceeds on links, which is where a site with no path for a crawler to follow pays twice.

Status returnedSitesExamples
40324congress.gov, oecd.org, imf.org, tripadvisor.com, temu.com, shein.com, revolut.com, lufthansa.com
4046nyp.org, nursingworld.org, thejakartapost.com, vietnamairlines.com, vlada.cz, etnhvac.com
5033washingtonpost.com, tokopedia.com, clinicbarcelona.org
5001santander.com
4011heb.com
4061who.int
4291cloudsmith.com
2021connectbase.com, a success code carrying no sitemap
timeout1spectator.co.uk, no response in 45 seconds
The 39 declared sitemaps that did not deliver a body on 1 October 2026, grouped by the status returned to LantadBot/1.0. The five non-absolute declarations are excluded because no request was made for them.

The refusal is not a decision about the crawler

The obvious reading of 24 refusals is bot defence, and the obvious remedy is to stop announcing yourself. That reading is wrong on this sample, and testing it is cheap. Every one of the 39 unreachable sitemaps was requested again on 1 October 2026 with two further user agent strings: a current desktop Chrome string, and the Googlebot string Google publishes. 37 of the 39 returned the identical status to all three requests. A declared sitemap that refuses LantadBot refuses a browser and refuses Googlebot too, so on those 37 sites this is a broken path or a blanket rule rather than a judgement about who is asking.

Two sites behaved differently, and they differ from each other. heb.com answered 401 to LantadBot and to the Googlebot string, and 200 to the Chrome string, which is a session or cookie gate rather than a crawler policy. laredoute.fr answered 403 to LantadBot and to the Chrome string, and 200 to the Googlebot string, which is the one case in the sample of a sitemap served on the strength of a claimed identity. That is worth naming precisely rather than dramatising: a user agent string is a request header anybody can send, so what laredoute.fr has built is a sitemap available to anything willing to call itself Googlebot, which is a different thing from a sitemap available to Google.

This matters for the question the product is actually asked, which is whether the crawlers behind AI answers can read a site. A refusal that applies to everyone is a site reliability problem with an SEO consequence. A refusal keyed to the user agent is an access policy, and policies keyed to a token are the ones that go stale, as this blog found when 1,004 robots.txt files named 2,209 tokens and 120 of them named a token Anthropic no longer documents. The 15 tokens Lantad evaluates are listed on the AI crawlers reference, and none of them appears in any of the 39 refusals as a cause, because the refusals did not depend on the token at all. Whether a crawler reaches robots.txt in the first place is a prior question, and one measured separately when 92 of 1,056 sites refused GPTBot the file they served to a browser.

heb.com/sitemap/siteindex.xml

  • LantadBot/1.0 401
  • Googlebot/2.1 401
  • Chrome 141 desktop 200
  • Reads as a session gate, not a bot rule

laredoute.fr/sitemap.xml

  • LantadBot/1.0 403
  • Chrome 141 desktop 403
  • Googlebot/2.1 200
  • Served to anything claiming to be Googlebot
The two of 39 unreachable sitemaps whose status depended on the user agent string, requested three times each on 1 October 2026. The other 37 returned the same status to all three.

Nobody broke the 50,000 URL limit

The 827 sitemaps that arrived held 992,251 loc elements between them. Not one file exceeded either published ceiling, and the margins are wide enough that the ceilings are not the constraint anybody should be checking first. The largest single urlset held exactly 50,000 URLs, at regeringen.se, which is a site paginating correctly rather than a site in breach. The next largest were nato.int at 47,832 and pa.gov at 42,977. The median urlset in the sample held 297 URLs.

The byte ceiling is further away still. The largest uncompressed body was 23,544,186 bytes, at salesforce.com, which is 44.9 percent of the 52,428,800 the protocol allows. ovhcloud.com came second at 15,943,455 bytes. The median body across all 827 was 4,568 bytes. Sitemap index files behave the same way: the largest in the sample was nature.com with 40,321 children, 80.6 percent of the 50,000 an index may list, and the median index named nine. So on this corpus the two numbers a sitemap validator is usually built around produced zero findings in 827 files, and the four checks after them produced all of the findings.

That is worth stating plainly because it is inconvenient for the category rather than for anybody's site. A size check is the easiest thing to automate and the easiest thing to put in a feature list, and on a thousand real sites it is close to dead weight. The checks that fired were about delivery and about whether an entry resolves, neither of which is a size question. The same pattern showed up in the robots.txt validator run, where every response over the 500 kibibyte parsing limit turned out to be an HTML page rather than a large rule set, so the size rule was really a second detector for a different mistake.

Nine of the 827 bodies arrived gzip compressed, among them jstor.org, airbnb.com, geisinger.org and paris.fr, and all nine decompressed cleanly and are counted at their uncompressed size, which is the size the protocol's limit is written against. Fifty four were served with a content type of application/rss+xml, and 53 of those 54 turned out to contain a urlset rather than a feed, so the content type is mislabelled rather than the file being a feed. Neither document requires a particular content type for a sitemap, so neither case is counted as a defect here. What a sitemap cannot tell a crawler, in any of these files, is that a page changed: that claim lives in lastmod, and 340 of 440 pages disagreed with the Last-Modified header their own server sent when this blog checked on 29 September. Discovery and freshness are separate problems, and only the first is what a sitemap is good at, which is also why a sitemap is not the lever for getting cited in AI Overviews.

  • Largest urlset, regeringen.se, 50,000 of 50,000 URLs 100% of limit At the ceiling, not over it
  • Largest index, nature.com, 40,321 of 50,000 children 81% of limit 80.6 percent of the limit
  • Largest body, salesforce.com, 23,544,186 of 52,428,800 bytes 45% of limit 44.9 percent of the limit
  • Longest loc, 1,080 of 2,048 characters 53% of limit 52.7 percent of the limit
  • Median urlset, 297 of 50,000 URLs 0.6 The typical file is nowhere near
How close the largest file in each category came to its published ceiling, measured on 1 October 2026 across the 827 sitemaps that delivered a body. Values are percentages of the limit the sitemaps.org protocol sets.

Where the 827 files that arrived actually failed

Six of the 827 bodies contained neither a urlset nor a sitemapindex element, and four of those six were HTML: unesco.org, walmart.com, regions.com and daytona.io each answered their declared sitemap URL with a web page and an HTTP 200. That is the sitemap version of the failure this blog measured as a soft 404 on 125 of 1,094 sites, and it is the worst of the format failures because the status code says the request succeeded. The other two, dbs.com and visitdubai.com, served XML with no sitemap root element at all, dbs.com at 977,154 bytes of it. A seventh case is insee.fr, which served a syntactically valid sitemapindex of 124 bytes containing no children, which parses perfectly and names nothing.

Fifteen of the 827 omitted the sitemaps.org 0.9 namespace declaration on an otherwise valid sitemap, among them cloudflare.com, marriott.com, kayak.com, xero.com, affirm.com, mcgill.ca and rte.ie. Every one of them parses in practice, because parsers in the wild match on element names rather than on the namespace, so this is the mildest finding in the post and is reported at that weight. It is listed because it is the kind of thing a strict validator reports as an error and a crawler ignores, and a site owner deserves to know which of those two they are looking at.

Seven sites listed at least one loc value that is not an absolute URL, 312 entries between them. On four of the seven every entry is relative and the file is therefore unusable as written: uio.no lists 18 of 18 as relative paths, elderlawgroupwa.com 9 of 9, soulyrested.com 5 of 5, and fornidental.com and villagedentaldtc.com 2 of 2 each. On the other two the fault is partial and harder to notice, which makes it more durable: oliverbonas.com lists 268 relative entries among 6,661, and rcgp.org.uk 8 among 1,586. Google's instruction that it crawls URLs "exactly as listed" is the operative sentence, and a relative path listed exactly as written is not a page. No file in the sample broke the 2,048 character limit; the longest loc was 1,080 characters.

Forty two sites listed at least one loc on a host other than the host that served the sitemap. On 37 of those the two hosts share a registrable domain, the ordinary apex against www split or a CDN subdomain, which is a scope question rather than a mistake. On five the sitemap is served from a different registrable domain entirely: santafe.edu declares one on s3.amazonaws.com, aboutyou.de and bettersheets.co on storage.googleapis.com, bradesco.com.br on banco.bradesco and sbi.co.in on sbi.bank.in. Under the protocol's location rule those files sit outside the scope of the URLs they list, and Google writes the consequence as a conditional: "Unless you submit your sitemap through Search Console, a sitemap affects only descendants of the parent directory." To see what a crawler takes from a single page rather than from an inventory, the what GPTBot sees tool fetches one URL the way a crawler would.

FindingSitesShareNamed examples
At least one loc on a different registrable domain50.6%santafe.edu, aboutyou.de, sbi.co.in
No urlset and no sitemapindex element60.7%unesco.org, walmart.com, dbs.com
HTML served at the sitemap URL with a 20040.5%unesco.org, regions.com, daytona.io
At least one relative loc value70.8%uio.no, oliverbonas.com, rcgp.org.uk
Zero loc elements in a valid document70.8%insee.fr, at 124 bytes
At least one loc cross host, same registrable domain374.5%paypal.com, salesforce.com, coolblue.nl
Missing the sitemaps.org 0.9 namespace151.8%cloudflare.com, marriott.com, kayak.com
Format findings across the 827 sitemaps that delivered an HTTP 200 body on 1 October 2026. Shares are of 827. The namespace row is listed last because every file in it parses in practice.

451 sitemap indexes, and whether the children answer

A sitemap index is a promise deferred, so validating one means following it. 452 of the 827 bodies were sitemapindex documents rather than urlset documents, which is the majority of the sample and the reason a validator that stops at the first file reports almost nothing useful about large sites. 451 of the 452 listed at least one child. The first child of each was fetched on 1 October 2026, which recovered a child URL on 445 of them.

436 of those 445 first children answered HTTP 200. Nine did not: justice.gov, timesofisrael.com, box.com, stuff.co.nz, decathlon.fr and boots.com each answered 403 for a child their own index names, dtu.dk answered 503, and tiki.vn and meininger-hotels.com answered 404. The meininger-hotels.com entry carries an escaped ampersand in its query string, which looked like a plausible cause until both the escaped and the unescaped form were requested and both returned 404, so it is counted as a genuine missing child. Twenty one of the 436 children were themselves another sitemapindex, so the inventory is two levels deeper than the robots.txt line suggests, and the median child urlset held 235 URLs.

Nine of 445 is a small number and it is reported as one. The point is not the rate, it is that the failure is one level below where anyone looks: the index at the address in robots.txt answered 200, a validator checking that address reports a pass, and a slice of the site's URL inventory is still unreachable. That is the same structural problem as 985 links found by crawling against 200,712 URLs declared in sitemaps, where the number a tool reports depends entirely on how far it walked.

It is also the context for the file that is supposed to make all of this unnecessary. An llms.txt is a hand curated list, and when this blog compared the two on the 239 sites publishing both it found a median of 23 URLs against 969 in the sitemap. The sitemap is the complete inventory and the thing worth repairing; the curated file is a summary of it. A site whose sitemap 403s has neither, and a site whose index leads to a 403 has part of one. Sites that publish a feed instead are a third case, measured when 86 of 762 declared an RSS feed.

The path a crawler walks from robots.txt to a URL, with the count that survived each step on 1 October 2026. Counts narrow at every edge, and the two refusal steps are where the inventory is lost.

What this does not measure

This run measured what a site publishes, not what any crawler did with it. No AI crawler's behaviour is observed here at all. The nearest evidence on that question is a reading of the vendors' own pages rather than of their traffic: the crawler documentation published by OpenAI, Anthropic and Perplexity, read by Lantad on 8 August 2026, mentions no sitemap on any of the three pages while all three describe robots.txt handling in detail. So a repaired sitemap is a defensible thing to want and an undocumented thing to rely on, and this post does not claim that fixing one of these 44 sites would produce a citation.

Three limits in the method are worth naming rather than leaving for a reader to find. Only the first Sitemap value in each robots.txt was fetched, and 219 of the 871 files declared more than one, booking.com listing 434 of them, so a site whose first sitemap is broken may well serve a working second. Only the first child of each index was followed, so the 436 figure is a probe and not a survey of every child. And the host scope check compares the loc host against the host that served the sitemap, which is the coarse version of the protocol's rule: the real rule is scoped by path, so a file at example.com/catalog/sitemap.xml listing example.com/images/ URLs is out of scope by the specification and counted as compliant here.

What the run does establish is narrow and solid. On 1 October 2026, on this corpus, the sitemap a site names in its own robots.txt failed to produce a document on 44 of 871 sites, the refusal was identical for a browser and for Googlebot on 37 of the 39 that were refused, and the two size ceilings a sitemap validator is usually built to check were not reached by any of the 827 files that did arrive. A single site's own numbers are the only ones that matter to that site, and the research page carries the aggregate picture across everything scanned rather than this one corpus.

The broader point is about where to spend an hour. Generative engine optimization attracts a lot of advice about new files to publish, and the measurement here keeps landing on an older and duller finding: the files these sites already publish, and already point at, often do not arrive. AI visibility begins with a document that answers when it is asked for, and 44 sites in this sample are failing that test while their robots.txt advertises otherwise.

  • Delivery of a declared sitemap Measured One GET for the first Sitemap value on each of 871 hosts, status recorded.
  • User agent sensitivity of a refusal Measured All 39 unreachable sitemaps re-requested with a Chrome and a Googlebot string.
  • Format compliance of the files that arrived Measured 827 bodies checked against the sitemaps.org protocol and Google's documentation.
  • Every sitemap a site publishes Not measured Only the first Sitemap value was fetched; 219 sites declared more than one.
  • Every child of every index Not measured One child per index, so 436 is a probe rather than a census.
  • Path scope under the protocol's own rule Partly measured Host compared, not directory, so some out of scope entries count as compliant.
  • What any AI crawler did with these files Not measured No crawler traffic is observed here, and three vendor docs name no sitemap.
What this run can and cannot support, stated before anybody quotes a figure from it. Measured 1 October 2026 on the 1,419 host corpus committed to this repository.

Written by

Lantad

Published .

A sitemap validator is usually sold against the two numbers everybody can recite: 50,000 URLs and 50 megabytes. Those are the ceilings the protocol writes down, so they are the ceilings the tools check. This run checked them against 827 real files and not one file was anywhere near either. What went wrong instead was more basic, and no size check would have caught it: on 44 sites the sitemap named in robots.txt did not produce a document at all.

Common questions

What does a sitemap validator check?

Whether the file arrives and whether its contents parse. Delivery comes first: the value in the robots.txt Sitemap line has to be an absolute URL and the request for it has to answer HTTP 200. Format comes second: a urlset, sitemapindex, feed or URL list, under 50,000 entries and 52,428,800 bytes, with every loc a fully qualified absolute URL under 2,048 characters and inside the sitemap's own location. On this corpus the delivery checks produced 44 findings and the two size checks produced none.

Is it a problem if my sitemap returns 403?

Yes, and it is the most consequential finding in this run. 24 of 871 declared sitemaps answered 403 on 1 October 2026. Neither the sitemaps.org protocol nor Google's documentation defines a fallback for a refused sitemap, so a crawler gets no URL inventory and discovers pages by following links instead. Checking it takes one request, and the refusal is usually not crawler specific: 37 of the 39 unreachable sitemaps returned the same status to a desktop Chrome string and to the Googlebot string as to LantadBot.

How many URLs can one sitemap hold?

50,000, and no more than 52,428,800 bytes uncompressed, which the sitemaps.org protocol and Google's sitemap documentation both state. A sitemap index may list 50,000 sitemaps under the same byte limit. These limits were not the binding constraint on this sample: across 827 files holding 992,251 loc elements, the largest urlset held exactly 50,000 URLs, the largest body reached 44.9 percent of the byte limit, and the largest index named 40,321 children.

Do AI crawlers use sitemaps?

Their own documentation does not say so. Lantad read the crawler documentation published by OpenAI, Anthropic and Perplexity on 8 August 2026 and none of the three pages mentions a sitemap, while all three describe robots.txt handling in detail. That is an argument for not resting a visibility claim on a sitemap, and not an argument for leaving a broken one in place: a sitemap that answers 403 also tells a conventional search crawler nothing, and robots.txt is advertising it either way.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.