BlogFindings
Article schema: 27 of 165 article pages declared a headline the page's own h1 does not contain
Lantad followed the home page's own links one level deeper on the 1,027 hostnames in this repository's committed industry corpus on 27 September 2026 and read 1,178 interior pages with no JavaScript executed. 176 of them declared exactly one Article node in JSON-LD. On 27 of the 165 that also served an h1, neither string contained the other, and 30 of the 176 declared no dateModified at all.
This run checked them mechanically. Lantad requested /robots.txt from all 1,027 hostnames in the committed industry corpus, evaluated each interior path for LantadBot before asking for it, then followed the home page's own links one level deeper and read up to two same-site interior pages per host as raw bytes with no JavaScript executed. 1,178 interior pages answered HTTP 200 with HTML across 611 hosts. 586 carried at least one JSON-LD block that parsed. 194 of those declared at least one node typed as an Article, a NewsArticle or a BlogPosting, and 176 declared exactly one such node, which is the denominator used for everything below: on a page carrying one Article node there is no question about which node describes the page.
In short
- Article schema is the one Google feature guide that requires nothing: the Article documentation, carrying Last updated 2026-09-08 UTC and read on 27 September 2026, states that there are no required properties and lists all five of author, dateModified, datePublished, headline and image as recommended only.
- Measured on 27 September 2026, 113 of the 176 interior pages declaring exactly one Article node carried all five recommended properties, and 3 carried none of the five while still declaring themselves an Article.
- On 27 of the 165 pages that declared a headline and also served an h1, neither string contained the other after HTML entities were decoded: stripe.com declared a newsroom headline against an h1 reading Stripe logo, thejakartapost.com declared the article title against an h1 reading The Jakarta Post, and typeform.com declared a headline ending [2025] above an h1 that says 2026.
- dateModified was the scarcest of the five, absent on 30 of the 176 pages, and on 34 of the 145 that declared both dates it was a byte-identical copy of datePublished, so the field that carries the freshness signal repeated the one beside it.
- 13 date values across four hosts were not the ISO 8601 that Google's Article documentation asks for, among them 24/09/2026 09:27:34 on bhf.org.uk, 28-08-2026 13:00 CEST on santander.com and Sun Sep 27 2026 on mejuri.com, which is the day the page was fetched.
| Stage | Pages or hosts | What it excludes |
|---|---|---|
| Industry corpus seed file | 1,027 hosts | Nothing. One committed sampling frame. |
| robots.txt produced a status | 998 hosts | 29 where the request failed at the transport layer |
| Allowed LantadBot at the root | 936 hosts | 62 by an explicit rule or a 5xx treated as disallow |
| Home page answered 200 with HTML | 697 hosts | 239 that answered otherwise |
| Interior pages fetched 200 with HTML | 1,178 pages | 26 disallowed for that path, 49 non-HTML, 1 that threw |
| Carried parseable JSON-LD | 586 pages | 591 with no block, 1 whose only blocks failed to parse |
| Declared an Article node | 194 pages | 392 that carried schema and no Article node |
| Declared exactly one Article node | 176 pages | 18 pages carrying 272 nodes between them |
What does article schema actually require?
Nothing, and that is not a figure of speech. Google's Article documentation, carrying Last updated 2026-09-08 UTC when it was read on 27 September 2026, says it in one sentence: there are no required properties; instead, add the properties that apply to your content. It then lists five properties as recommended, being author, dateModified, datePublished, headline and image, and one constraint on the type itself: Article objects must be based on one of the following schema.org types, Article, NewsArticle, BlogPosting.
That is a much weaker contract than the neighbouring feature guides, and the difference is the whole reason this measurement is worth taking. A rating has to carry a ratingValue or the review snippet earns nothing. A breadcrumb has to carry more than one item to be a trail rather than a single link. An Article node satisfies its documentation by existing. So a validator passing is not evidence about an Article node the way it is evidence about the others, and the only question left is an empirical one: given that nothing is required, what do real pages put in?
The vocabulary underneath is looser again. schema.org/Article, read on 27 September 2026, defines the type as an article, such as a news article or piece of investigative report, and reports its usage as 10M+ Domains from monthly aggregations of Google's web index for August 2026. schema.org/datePublished is defined as the date of first publication or broadcast, which is unambiguous. schema.org/dateModified is not: it is defined as the date on which the CreativeWork was most recently modified or when the item's entry was modified within a DataFeed. That second clause is a legitimate reading in which a feed rebuild is a modification, and it is worth holding in mind before treating any dateModified value as a statement about the text.
Two pages in the corpus used a subtype rather than one of the three types Google names. eltiempo.com declared a ReportageNewsArticle on one page and a BackgroundNewsArticle on another. Both sit under NewsArticle in the schema.org hierarchy, which reads Thing then CreativeWork then Article then NewsArticle then ReportageNewsArticle, so both are based on a supported type and neither is a defect. It is a detail worth stating because a naive string comparison against the three names Google lists would report it as one, which is the same class of mistake as reading a JSON-LD syntax defect out of a file that parsed.
| Property | Status | What the documentation says |
|---|---|---|
| author | Recommended | Person or Organization. Include all authors in the markup. |
| datePublished | Recommended | The date and time the article was first published, in ISO 8601 format. |
| dateModified | Recommended | Most recently modified, ISO 8601. The Rich Results Test shows no warning when it is absent. |
| headline | Recommended | The title of the article. Consider using a concise title. |
| image | Recommended | Representative of the article, rather than logos or captions. |
| Anything required | None | There are no required properties; add the properties that apply to your content. |
How many article pages carried all five recommended properties?
113 of the 176, which is 64.2 percent, and the distribution either side of that is narrow. 41 pages carried four of the five, 12 carried three, 3 carried two, 4 carried one and 3 carried none at all. So two thirds of the pages that bother to declare an Article declare the full recommended set, and the tail is thin rather than long. That is a better result than this corpus has returned for most markup properties measured on it, and a good deal better than the 313 of 615 home pages that put an @id on no schema node.
Ranked by presence, the five come out as datePublished on 168 pages, headline on 167, image on 154, dateModified on 146 and author on 140. The ordering is the interesting part. The two properties a content management system can fill without anybody deciding anything, being the publication timestamp and the title, are the two that are almost always there. The two that need an editorial decision, being who wrote it and whether it has been changed since, are the two that go missing, 36 and 30 times respectively. That is the same shape as the 95 of 111 article pages that declared an author on this corpus five days earlier, where the failure was not absence so much as 8 nodes typed as a Person that named a publication instead of a human being.
The three pages carrying none of the five are worth naming because of what they are rather than how many. worldbank.org declared a bare Article node on a results page, and dana-farber.org declared one on a departments and centers page and again on a clinical trials page. None of the three is an article. They are section pages that inherited an Article node from a template, and an Article node with no headline, no author, no image and no dates is a statement that something on this page is an article and nothing more. A machine reading it learns that a template was applied.
Strata sharpen this. Of the 48 news pages in the sample, 40 carried all five. Of the 16 education pages, 13 did. Government managed 3 of 15 and healthcare 4 of 18, and those two strata account for 8 of the 30 missing dateModified values between them. Publishers whose product is dated by nature declare dates; institutions whose pages are standing reference material largely do not, which is consistent with the finding that 250 of 647 home pages sent no machine-readable statement of change at all.
Does the declared headline match the heading on the page?
165 of the 176 pages declared a headline and also served an h1, so the two can be compared directly. Both strings were decoded for HTML entities, stripped of accents and punctuation and lowercased before comparison, because a raw comparison reports a difference every time a page serves the same words with a numeric entity in them, which several of the Swedish and Spanish pages here do. After that normalisation, 124 pages declared a headline identical to the h1. 14 more had one string contained inside the other, which is a site name suffix or a section prefix rather than a different claim: nasa.gov declared NASA Media Contacts - NASA above an h1 of NASA Media Contacts, and usps.com added a pipe and the brand.
27 pages were left in which neither string contained the other. That is 16.4 percent of the comparable pages, and the cases are not near misses. stripe.com declared, on two newsroom pages, headlines naming the Instant Checkout announcement and the NVIDIA collaboration, above an h1 reading Stripe logo, because the only h1 in the document is the alt text of the masthead image. thejakartapost.com declared the article title on two pages whose only h1 reads The Jakarta Post. rki.de declared the article title above an h1 reading Navigation und Service. frontiersin.org declared two different article headlines above the same h1, Frontiers | Science news. In every one of those the declared headline is the better string, and the h1 is a piece of furniture.
Two cases run the other way and are more instructive. typeform.com declared a headline ending in the bracketed year [2025] on a comparison page whose h1 says 2026: the visible heading was updated and the markup was not, so a machine that trusts the markup dates the page a year early. canonical.com declared a 344 character headline that is not a title at all but the opening two sentences of the post, beginning Developers now benefit from consistency and repeatability for cutting-edge workflows, above an h1 that is a proper title. Google's own guidance on that field is one line, that you should consider using a concise title as long titles may be truncated on some devices, and 11 of the 176 headlines ran past 110 characters.
11 pages served no h1 element at all, among them two worldbank.org press pages, two santander.com stories and westwing.de. On those the Article node is the only title-shaped string a crawler that executes nothing will find, which is a decent argument for filling it carefully and a poor argument for leaving the document without a heading: the 311 of 704 article pages that ran over 300 words with no heading measured on this corpus are the same problem seen from the other side.
stripe.com newsroom page
- headline: Stripe powers Instant Checkout in ChatGPT and releases Agentic Commerce Protocol
- h1: Stripe logo
- The only h1 in the document is the masthead image alt text.
typeform.com comparison page
- headline: Typeform vs. Jotform: Which is the best form builder? [2025]
- h1: Typeform vs. Jotform: which is the best form builder in 2026?
- The heading was updated for 2026. The markup still says 2025.
Why does dateModified so often say nothing?
145 of the 176 pages declared both dates, and on 139 of those both values resolved to a calendar day. 74 of the 139 gave the same day for both. 34 of the 145 went further and declared dateModified as a byte-identical copy of datePublished, character for character, including the fractional seconds: regeringen.se declared 2025-11-12T09:15:21.0000000+01:00 as both, nature.com declared 2026-09-25T00:00:00Z as both, and helios-gesundheit.de declared 2026-07-22T00:00:00Z as both.
A same-day pair is not by itself a defect. A news story published and corrected within the hour genuinely carries two same-day timestamps, and 19 of the pages here declared a dateModified on 27 September, the day of the crawl, because they were that day's news: cnn.com, aljazeera.com, variety.com, time.com and rte.ie among them. A byte-identical pair is a different object. It says the page has not been modified since publication, which is a real claim, but it is also exactly what a template emits when it is handed one timestamp and asked to fill two fields, and nothing in the markup distinguishes the two situations. This is the same trap as a sitemap lastmod value being an assertion rather than a measurement, and the honest reading of a byte-identical pair is that the field is uninformative rather than that the page is unchanged.
45 pages declared a gap of more than 30 days, and 23 of those a gap of more than a year, which is where dateModified starts doing real work. nationwidechildrens.org declared a publication date of 20 November 2013 and a modification date of 10 February 2026. gov.scot declared 29 October 2018 and 26 September 2026. foxnews.com declared 1 March 2011 and 6 February 2019. Those pages are telling a machine something it could not otherwise know and could not guess from the URL, which is the case for the property existing.
Four pages declared a dateModified earlier than their datePublished, which cannot be true of the same document. restofworld.org declared 2026-09-23T06:00:00-04:00 published against 2026-09-22T10:27:08-04:00 modified on one story and the same one day inversion on another, and typeform.com declared 2026-09-04T00:14:50.000Z published against 2026-09-03T19:22:44.624Z modified on two pages. The likely mechanism is an embargoed publication timestamp set in the future at the moment the file was last written, which is a plausible editorial workflow and still an impossible pair of statements. Google's guidance on publication dates, carrying Last updated 2025-12-10 UTC, says its systems look at several factors to determine a best estimate precisely because all factors can be prone to issues, and a dateModified before a datePublished is what that sentence is about.
Do the declared dates parse as ISO 8601?
Mostly. 161 of the 168 datePublished values and 140 of the 146 dateModified values matched ISO 8601, leaving 13 values across four hosts that did not. bhf.org.uk declared 24/09/2026 09:27:34 as both dates on one page and 21/09/2026 09:21:05 as both on another, in a day-first order that is unambiguous to a British reader and ambiguous to a parser. santander.com declared 28-08-2026 13:00 CEST and 22-09-2026 08:00 CEST, which is day-first with a named zone rather than an offset. zapier.com declared February 11, 2020 and April 20, 2023, which are prose dates in the markup. mejuri.com declared Sun Sep 27 2026, which is JavaScript's default date string and is also the day the page was fetched, so whatever it describes it is not a publication date.
Then there is the part nobody fails a validator on. 15 of the 161 well-formed datePublished values carried no timezone offset and 9 of the 140 dateModified values did not, and 10 datePublished values were a bare calendar date with no time in them at all. Google's Article documentation is explicit about the consequence: it recommends that you provide timezone information, otherwise it will default to the timezone used by Googlebot. For a daily news page that defaulting can move a timestamp across a day boundary, and the pages most likely to omit the offset are the institutional ones where the day is all that was ever known.
None of this is visible to a reader, which is what makes it worth measuring. A date in the wrong format still renders correctly on the page, because the page renders a different string: the one in the template, not the one in the JSON-LD. The markup is the only copy a client that does not execute JavaScript will read, and on this corpus that client sees 7.6 percent less prose than a browser does, so the two copies drifting apart is not a hypothetical. It is the same failure mode as 429 of 615 pages carrying parseable JSON-LD that named no language in it: correct markup, quietly missing the one field that resolves an ambiguity.
There is a second reason to care about the format specifically, which is that Google's publication dates guidance asks for something this measurement cannot see. It says to add a user-visible date to the page and feature it prominently, and to label it with text like Publish or Last updated. Whether each of these 176 pages does that is a judgement about layout, not a string comparison, and this run did not attempt it. What it can say is narrower and still useful: on 46 of the 161 pages whose datePublished resolved to a day, that calendar day appeared nowhere in the served text in any of the formats tried, which include ISO order, day first, month first and month names in the eight languages present in the corpus. Some of those 46 show a neighbouring day instead, as clickhouse.com does in declaring 22 April 2024 and displaying Apr 23, 2024, so 46 is an upper bound on pages showing no matching date rather than a count of pages showing no date.
| Host | Value as served | Why a parser struggles |
|---|---|---|
| bhf.org.uk | 24/09/2026 09:27:34 | Day first, no zone. Declared as both dates on the same page. |
| bhf.org.uk | 21/09/2026 09:21:05 | Same form on a second page, again as both dates. |
| santander.com | 28-08-2026 13:00 CEST | Day first with a named zone rather than an offset. |
| santander.com | 22-09-2026 08:00 CEST | Second page, same form. |
| zapier.com | February 11, 2020 | A prose date in a field specified as ISO 8601. |
| zapier.com | April 20, 2023 | Second page, same form. |
| mejuri.com | Sun Sep 27 2026 | JavaScript's default date string, and the day of the crawl. |
What this run did not measure
It did not measure what any AI crawler does with these fields. No claim above says that a missing dateModified costs a citation, that a headline disagreeing with an h1 changes which string an answer engine quotes, or that any of the 27 pages is ranked differently for it. Lantad reads what a page hands a client that executes nothing, which is the layer it can measure; what happens inside a model is not that layer, and the vendors whose AI crawler documentation would settle it do not address extraction at all. Treating Google's stated preferences as a proxy for every engine's behaviour would be an inference dressed as a finding.
The sampling has four limits worth stating plainly. Interior pages were chosen by the shape of their URL, preferring paths containing blog, news, article, press and the like, which biases the set towards pages that are articles and says nothing about the rest of a site. Up to two pages per host means a large publisher and a small one contribute the same weight, so no percentage here is a percentage of the web. Every request was one attempt from one network location with a twenty second timeout, so a host that was briefly slow is absent rather than counted. And the 176 page denominator deliberately drops the 18 pages that declared more than one Article node, including a bostonglobe.com page carrying 73 and an irishtimes.com page carrying 45, because on those the question of which node describes the page has no clean answer. 56 of the 60 irishtimes.com nodes carried neither a headline nor a name, which is one site's recirculation module rather than a finding about anything.
What the run does establish is narrow and checkable. Given a feature guide that requires nothing, 64.2 percent of pages declaring an Article still filled all five recommended properties, and the properties that went missing were the two that need a human decision rather than a template variable. The disagreements that showed up were not between the markup and a specification; they were between the markup and the page it sits in, which is the class of defect no validator reports because both halves are individually valid.
The method is written down. How Lantad scores what it reads covers the extraction path and the research page covers the corpus and the earlier runs against it, including the measurements this post compares itself to. Anybody can repeat this one on a single page in a minute: fetch it with curl, read the JSON-LD, and check the headline against your own h1 and the dates against what your page says. Two of the three fields most worth getting right in structured data are strings you can verify by eye, and the entity confidence a machine builds from them is only as good as the weakest of the two copies you publish.
Flow: Industry seed file, 1,027 hosts to GET /robots.txt as LantadBot; GET /robots.txt as LantadBot to Parse with core/src/robots.ts; Parse with core/src/robots.ts (allowed) to GET / on 936 allowed hosts; GET / on 936 allowed hosts (200 and HTML) to Read same-site links from the HTML; Read same-site links from the HTML to GET up to 2 interior pages per host; GET up to 2 interior pages per host to Parse every ld+json block; Parse every ld+json block to 176 pages with one Article node.
Lantad
Published .
Most of the structured data a page ships is a description of a thing. This is an organisation, this is its name, this is a product and here is its price. An Article node is different in one respect that matters for an answer engine: it claims to say what the page is called and when it was written. Those two facts are exactly what a machine needs in order to quote the page and to decide whether the quote is still current, and they are the two facts a reader can check against the page in front of them in about four seconds.
Common questions
What does article schema require?
Nothing. Google's Article structured data documentation, carrying Last updated 2026-09-08 UTC and read on 27 September 2026, states that there are no required properties and that you should add the properties that apply to your content. It lists author, dateModified, datePublished, headline and image as recommended, and it does constrain the type: an Article object must be based on Article, NewsArticle or BlogPosting, which includes subtypes of those such as ReportageNewsArticle. The practical consequence is that a validator passing tells you almost nothing about an Article node, so the only useful question is what the node actually carries.
Should the headline in my Article markup match my h1?
Google does not require it, and this measurement found 27 of 165 pages where it did not. The reason to make them agree is that they are two answers to the same question served to the same client, and a machine has no way to prefer one. Where they differed here, the declared headline was usually the better string, because the only h1 in the document was a masthead image alt text or a site name. The case that should worry you is the opposite one: typeform.com serves an h1 saying 2026 above markup saying 2025, so the copy a crawler reads is a year out of date.
Does a missing dateModified hurt AI visibility?
This run did not measure that and no figure here supports a yes or a no. What it measured is that 30 of 176 pages omitted the property and that 34 of the 145 declaring both dates made dateModified a byte-identical copy of datePublished, which means the field carried no information on roughly a quarter of the pages that had it. Google's publication dates guidance says its systems use several factors to estimate when a page was published or significantly updated, so the markup is one input among several rather than the deciding one.
How do I check my own article schema?
Fetch the page as bytes rather than in a browser, since the markup a crawler reads is the markup in the response. Find the application/ld+json blocks, parse them, and check four things: that the node type is Article, NewsArticle or BlogPosting, that headline says the same thing as your h1, that datePublished and dateModified are ISO 8601 with a timezone offset, and that dateModified is not simply a copy of datePublished. The last two are where this corpus failed most often and neither shows up in a rendered page.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.