BlogFindings
Sitemap lastmod: 340 of 440 pages disagreed with the Last-Modified header their own server sent
Lantad requested /robots.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026 as LantadBot. 859 declared a sitemap and 785 of those resolved to a URL set holding 2,768,047 entries, of which 2,004,336 carried a lastmod. Two of those URLs per site were then fetched, and on the 440 responses that carried a Last-Modified header, 340 disagreed with the sitemap's claim about the same page by more than a day.
Lantad requested /robots.txt once from each of the 1,419 hostnames in the two corpus seed files committed to this repository, on 29 September 2026, as LantadBot, following redirects and executing no JavaScript. Where a file declared a sitemap, the first declared address was fetched, indexes were resolved one or two levels down to the first child holding a real URL set, and every lastmod in that file was read. Then, for each site, two of the URLs the sitemap had dated were requested directly, so the claim could be set against the one other timestamp the same server publishes about the same page. The two agreed on 100 of the 440 pages where both existed.
In short
- Lantad requested /robots.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026 as LantadBot/1.0: 859 declared a sitemap, 785 resolved to a URL set of at least two entries, and those 785 files held 2,768,047 URL entries of which 2,004,336 carried a lastmod.
- A sitemap lastmod is not verifiable on most sites that publish one: of 440 page fetches on 29 September 2026 that returned HTTP 200 with a Last-Modified header, 340 disagreed with the sitemap's claim about the same URL by more than 24 hours, and on 333 of those 340 the header was the later of the two.
- 94 of the 524 sitemap files read in full on 29 September 2026 gave every one of their URLs an identical value, covering 50,646 URLs, and the sitemap protocol says the date must be the date the page was last modified rather than the date the sitemap was generated.
- 182 of the 785 sitemaps read on 29 September 2026 carried no lastmod on any URL, and the publishing platform decides that more than the sector does: all 54 Wix and Squarespace sites, all 33 Shopify stores and all 26 WordPress sites emitted one, against 6 of 26 Framer sites.
- Google's sitemap documentation, carrying Last updated 2026-07-08 UTC, says it uses the lastmod value if it is consistently and verifiably accurate, and on this corpus the one other HTTP signal a crawler could check it against agreed on 100 of 440 pages.
| Step | Count | Note |
|---|---|---|
| Hostnames requested | 1,419 | 392 platform frame, 1,027 industry frame |
| robots.txt answered HTTP 200 with a body | 1,093 | 44 of those were an HTML document, not a text file |
| Declared at least one sitemap address | 859 | 3,118 Sitemap lines between them |
| Resolved to a URL set of two or more entries | 785 | 426 of the 785 reached through a sitemap index |
| URL entries read | 2,768,047 | Median file held 360 entries, largest held 50,000 |
| Entries carrying a lastmod | 2,004,336 | 72.4 percent of the entries read |
| Files carrying a lastmod on every URL | 484 of 785 | 119 more carried one on some URLs, 182 on none |
| Pages fetched to check a claim | 1,167 | Up to two per site, all returning HTTP 200 |
| Fetches returning a Last-Modified header | 440 of 1,167 | On 229 sites, so this test covers a minority |
| Claims the header contradicted by over a day | 340 of 440 | 333 of the 340 had the header as the later date |
Who publishes a sitemap lastmod, and what is it supposed to say?
859 of the 1,419 hostnames declared at least one sitemap address in a robots.txt that answered as a text file, which is consistent with the 839 of 1,016 files that declared one counted on a smaller frame earlier this month. 785 of the 859 then resolved to a URL set holding two or more entries. The 74 that did not are the usual mixture of a declared address returning 404, an index whose children were all empty, and a file whose bytes were not XML at all, which is the same class of defect counted when 232 of 1,059 robots.txt files carried a defect a validator can name.
Across the 785 files, 2,768,047 URL entries were read and 2,004,336 of them carried a lastmod. Counted by file rather than by URL, 484 put a lastmod on every entry, 119 on some entries but not all, and 182 on none at all. The partial group matters more than its size suggests, because the sitemap protocol at sitemaps.org makes lastmod optional per URL, so a file that dates half its pages is conformant and still gives a crawler no way to tell an undated page from an unchanged one.
The protocol also says what the value is supposed to mean, and it is unambiguous. The page at sitemaps.org/protocol.html states that lastmod is "the date of last modification of the page", that the date "should be in W3C Datetime format", and, in a sentence worth reading twice, that "the date must be set to the date the linked page was last modified, not when the sitemap is generated". The format it points at is the W3C Datetime note, which permits a bare date and requires a time zone designator whenever a time is present.
Conformance to the format is close to universal here and the exceptions are concentrated. Of 795,540 values examined, being every value in the 524 files read in full plus the first 5,000 in each of the 79 larger ones, 5,685 failed the format on four sites: one wrote the American month-day-year order, one used a space where the standard requires the letter T, and two omitted the time zone while supplying a time. 281,962 values, 35.4 percent of those examined, were a bare date with no time, which the protocol expressly allows and which costs a crawler nothing except resolution. Two values on two sites were dated more than a day into the future.
-
Valid W3C Datetime789,855 of 795,540 Parses under the format the sitemap protocol names, so a crawler can read it. -
Bare date, no time281,962 Expressly permitted. Costs resolution, not validity. -
Wrong format5,685 on 4 sites Month-day-year order, a space instead of T, or a time with no time zone. -
Dated in the future2 on 2 sites More than one day ahead of the request. A date a page cannot have reached. -
Files dating every URL484 of 785 The only shape from which a crawler can infer anything about an undated page. -
Files dating no URL182 of 785 Conformant, and it leaves the sitemap saying nothing about time.
Can a crawler check the claim against anything else?
Google's own sitemap documentation, at Google Search Central on building a sitemap and carrying Last updated 2026-07-08 UTC, sets a condition rather than an instruction: it says Google uses the lastmod value "if it's consistently and verifiably (for example by comparing to the last modification of the page) accurate", and that it ignores changefreq and priority outright. This site has already read that page and reported that a sitemap tells a crawler where pages are, not that they changed. What that post could not say, because it was reading documentation rather than sites, is how often the condition is met.
There is exactly one other timestamp about the same page that the same server publishes on its own. RFC 9110 defines Last-Modified in section 8.8.2 as a field giving "the date and time at which the origin server believes the selected representation was last modified", and the HTTP semantics specification adds, in the generation rules just below, that how that value is determined for any given resource is an implementation detail outside the specification's scope. So the header is not an authority either. It is simply the second opinion, and two independent signals that agree are worth more than one signal that cannot be checked.
For each of the 603 sites whose sitemap dated at least one URL, up to two of those dated URLs were requested directly as LantadBot. 1,167 of the 1,206 requests returned HTTP 200. 440 of those 1,167 responses, on 229 sites, carried a Last-Modified header at all, which is the first finding and an uncomfortable one: on the other 727 there is nothing to check the sitemap against, and a crawler wanting to confirm a date would have to download the page and diff it against a copy it may not hold.
On the 440 where both values existed, 100 agreed within 24 hours and 340 did not. The direction is almost uniform: on 333 of the 340 disagreements the header was later than the sitemap's claim, often by weeks, and on several the header moved between two requests seconds apart. That pattern is what a generated timestamp looks like rather than an edit record, and it is why this section states a disagreement rather than an error. The sitemap protocol anticipates exactly this, noting that lastmod "is separate from the If-Modified-Since (304) header the server can return, and search engines may use the information from both sources differently". Counted by site rather than by page, 162 of the 229 had every pair disagree and 38 had every pair agree.
Header contradicted the sitemap: 340
- Both timestamps present, more than 24 hours apart
- 333 of the 340 had the header as the later date
- 162 of 229 sites had every checked pair disagree
- Several headers moved between requests seconds apart
- Consistent with a generated or cache-fill timestamp
- Not a protocol violation, and not verification either
Header agreed with the sitemap: 100
- The two independent signals fall within a day
- 38 of 229 sites had every checked pair agree
- 29 more sites agreed on one page and not the other
- This is the shape Google's condition describes
- It is 8.6 percent of the 1,167 pages fetched
- The rest published a date nothing else confirms
What a build timestamp looks like when it reaches a sitemap
524 of the 603 sitemaps carrying a lastmod were small enough to be read in full, meaning every value was examined rather than the first 5,000. On 94 of those 524, every single URL carried the identical value. Those 94 files cover 50,646 URLs between them and the median one dates 116 pages to the same instant or the same day. A site does not edit 116 pages simultaneously. What produces that shape is a generator stamping the moment the file was written onto every row, which is the one thing the sitemap protocol names and rules out.
Widen the test from an identical instant to an identical calendar day and the count rises to 112 of the 524, so roughly one file in five carries no differentiation between its pages at all. The remaining files do differentiate, and there the distribution is the more useful number. Taking the freshest lastmod in each of the 600 files where at least one value parsed: 240 were dated within the last day, 134 within the last week, 82 within the last 30 days, 93 within the last year, and 51 carried nothing newer than a year old.
Those two tails describe opposite failures and both cost the same thing. A file whose newest date is a year old is either a site that genuinely has not changed or a generator that stopped running, and a crawler cannot tell which. A file that redates every page every night, which the 240 within-a-day group will contain some of, tells a crawler that everything changed, which is the same information as telling it nothing changed. Between them sits the population the signal was designed for, and nothing in the file marks which group a site belongs to.
This is the same failure mode measured from a different direction when 250 of 647 home pages sent no signal that anything had changed, and it is why a change signal is worth auditing rather than assuming. It also bears on how much weight the discovery layer deserves at all: this site recently found that on the 239 sites publishing both, an llms.txt named a median of 23 URLs against 969 in the sitemap, so the sitemap remains by far the larger surface, and the larger surface is the one carrying the unverifiable dates.
Which platforms emit a date, and which leave the element out
The 182 files carrying no lastmod at all are not spread evenly, and the variable that predicts them is the publishing platform rather than the industry. Every one of the 54 Wix and Squarespace sites in the platform frame emitted a lastmod, as did all 33 Shopify stores and all 26 WordPress sites. At the other end, 6 of 26 Framer sites did, and 7 of 14 static documentation sites. A site owner on one of those stacks has not made a decision about sitemap metadata. Their generator made it, and the Framer fix guide exists because that pattern repeats across every signal this scanner grades.
The industry frame varies far less: healthcare 46 of 53, news 47 of 57, finance 65 of 81, education 32 of 43, SaaS 78 of 107, travel 41 of 56, ecommerce 36 of 55 and government 28 of 39. The spread from the best sector to the worst is about 21 percentage points, against a spread from Wix to Framer of 77. Anyone drawing conclusions about an industry's technical maturity from a sitemap field is mostly measuring which website builders that industry buys.
One platform behaviour was visible only because it broke an earlier pass of this measurement, and it is worth recording. 448 of the 859 declared sitemaps were a sitemap index rather than a URL set, and 39 of those indexes listed a child sitemap named for agentic discovery, 32 of them in the Shopify stratum and 7 elsewhere in ecommerce. Each of those children held a single URL pointing at an agents.md file and carried no lastmod. Reading the first child of an index, which is the obvious implementation, therefore returns a one-line file on those sites rather than the product sitemap, and the first run of this measurement reported Shopify as the worst platform for lastmod coverage before the selection rule was changed to take the first child holding a real URL set. The corrected figure is 33 of 33. That a platform has begun shipping a discovery file aimed at agents is itself a change since this site counted one agent card across 1,419 sites ten days earlier.
| Platform stratum | Sitemaps read | Carried a lastmod | Share |
|---|---|---|---|
| Wix and Squarespace | 54 | 54 | 100% |
| Shopify | 33 | 33 | 100% |
| WordPress | 26 | 26 | 100% |
| SaaS marketing | 32 | 29 | 91% |
| Local media | 25 | 22 | 88% |
| SPA startups | 33 | 22 | 67% |
| Webflow | 33 | 20 | 61% |
| Bubble and no code | 18 | 11 | 61% |
| Static documentation | 14 | 7 | 50% |
| Framer | 26 | 6 | 23% |
What an AI crawler can do with a date it cannot verify
The crawler documentation published by OpenAI, Anthropic and Perplexity does not mention sitemaps at all, which this site read at source on 8 August 2026 and recorded when it first looked at this element. So the honest framing of every figure above is that it describes a signal aimed at classical search indexing that an AI crawler may or may not read. What is not in doubt is the mechanism underneath, because a crawler deciding whether to refetch a page has exactly three things available and the sitemap supplies the weakest of them.
The other two are HTTP validators, and this run counted those too. Of the 1,167 pages fetched, 528 carried an ETag and 440 carried a Last-Modified header, and 191 carried both, leaving 390 of the 1,167 with neither of the fields that would let a crawler ask "has this changed" without downloading the answer. 167 sent a Cache-Control header containing no-store, which instructs a client not to keep a copy at all and therefore removes the possibility of a conditional request on a later visit. A site in that group is telling every crawler to download the whole page, every time, and then telling it separately through a sitemap that the page was last edited in March.
The practical reading for a site owner is narrow and it is not "fix your lastmod". It is that a date nothing can confirm earns nothing, and that the two fields which can be confirmed are cheaper to get right because the server generates them. If you want a crawler to notice a change, change the bytes and let the validators move. If you want it to notice quickly, the discovery ping is a separate mechanism with its own limits, which this site covered when it found that IndexNow speeds discovery and not readability. And if the goal behind the question was crawl efficiency rather than freshness, the advice that circulates under that heading transfers to AI crawlers badly, as measured when crawl budget advice landed on AI crawlers and when no crawler vendor documented honouring Retry-After.
Freshness is also not the binding constraint for most sites in this corpus, which is worth saying plainly on a page about a freshness field. A sitemap that dates a page perfectly still only helps a crawler that can fetch and read the page, and readability is where the failures concentrate: see what GPTBot sees for the fetch, our robots.txt tester for the access rules, and structured data for what an engine extracts once it is in. A perfect lastmod on a page that returns no words is a precise timestamp on nothing.
Flow: Has this page changed? (2,004,336 dated) to Sitemap lastmod; Has this page changed? (440 of 1,167) to Last-Modified header; Has this page changed? (528 of 1,167) to ETag validator; Sitemap lastmod to Cross-check the two; Last-Modified header to Cross-check the two; Cross-check the two to 340 of 440 disagreed.
What this run did not measure
No crawler was observed fetching any of these sitemaps. Nothing here supports a claim that an engine read a lastmod, weighted one, or recrawled because of one, and any sentence in that shape would be an invention. What was measured is what the files declare and whether a second signal from the same server agrees.
The verification test covers a minority of the corpus and the minority is not random. It rests on the 440 responses that carried a Last-Modified header, on 229 sites, and a server that emits that header is more likely to be serving files from disk than generating pages, so the group is biased toward the simpler stacks. Two URLs per site were checked, and they were the first two dated entries in the file, which on many sites are the home page and a top-level section rather than a representative page. A disagreement is also not proof that the sitemap is wrong: as the protocol itself says, the two fields are separate, and a Last-Modified that tracks a cache fill will disagree with a perfectly maintained lastmod. The finding is that the claim cannot be confirmed from outside, not that it is false.
Reading one sitemap file per site caps what the figures describe. 426 of the 785 arrived through an index, and only the first child holding a real URL set was read, so a site whose other children are dated differently is represented by one of them. 79 files exceeded the 5,000 value cap and were examined only to that depth. Every request was one attempt, on one day, from one network location, so a host that refused this network may serve another, and a file regenerated the next morning is not tracked. The methodology page states the general limits, and the crawlability study is where the running corpus figures live.
Finally, the corpus is an editorial sampling frame assembled for platform and industry coverage rather than a random draw of the web, so every rate here supports a statement about these 1,419 hostnames and nothing wider. That constraint applies to the generative engine optimization conclusions as much as to the counts, and where an earlier finding on this subject came from a smaller frame, this post says so rather than presenting one number as the web's. The nearest neighbour to that caution in this run is the SearchAction measurement, where 57 of 224 declared search endpoints pointed at a path the same site's robots.txt disallows: the same pattern of a declaration that nothing checks.
- Sitemap files read and parsed 785 files, 2,768,047 URL entries, one file per host.
- Claims cross-checked against a second signal 440 of 1,167 fetched pages carried a Last-Modified header.
- Whether any engine reads lastmod No crawler was observed and no server logs were held.
- Whether a disagreeing lastmod is wrong The protocol treats the two fields as separate on purpose.
- Whether a lastmod changed a ranking or a citation No ranking or answer engine output was measured.
- Coverage beyond the first sitemap per site 426 of 785 arrived via an index and one child was read.
Lantad
Published .
A sitemap is the cheapest discovery surface a site owns, and the lastmod element is the only part of it that claims anything about time. Every guide to preparing a site for answer engines recommends keeping a sitemap lastmod accurate, and almost none of them say what accurate would be measured against. This run went and measured it.
Common questions
What is sitemap lastmod supposed to contain?
The date the page itself was last modified. The sitemap protocol at sitemaps.org states the value should be in W3C Datetime format and that the date must be set to the date the linked page was last modified, not when the sitemap was generated. Lantad found 94 of 524 sitemap files read in full on 29 September 2026 giving every URL an identical value, which is the shape a generation timestamp produces.
Does Google use the lastmod value?
Conditionally. Google's sitemap documentation, carrying Last updated 2026-07-08 UTC, says it uses the lastmod value if it is consistently and verifiably accurate, giving comparison to the last modification of the page as the example, and says it ignores changefreq and priority. On this corpus the second signal available for that comparison agreed on 100 of 440 pages.
Do AI crawlers read sitemaps?
The crawler documentation published by the vendors behind the main citation crawlers does not mention sitemaps, which Lantad recorded when it read those pages on 8 August 2026. This run measured what sites declare, not what any engine consumes, so no figure here supports a claim that GPTBot, ClaudeBot or PerplexityBot acted on a lastmod.
If my sitemap has no lastmod, is that a problem?
It is not a protocol violation, since the element is optional per URL, and 182 of the 785 sitemaps Lantad read on 29 September 2026 carried none. What it removes is a change signal. The two HTTP validators, ETag and Last-Modified, are stronger because a crawler can test them, and 390 of the 1,167 pages fetched in this run carried neither.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.