BlogFindings

Speakable schema: 8 of 1,114 home pages declared it, and 5 of 12 resolvable selectors pointed at no readable text

Lantad requested the robots.txt and then the home page of all 1,419 hostnames in this repository's committed corpus on 30 September 2026, read the delivered bytes with no JavaScript executed, and analysed the 1,114 pages that were both allowed and readable. 8 of them declared the schema.org speakable property, and a second pass over the news strata found it on 5 of the 69 hostnames where an Article typed page could be identified. Of the 12 CSS selectors those declarations used that name a single element by id, class or tag name, 4 matched no element in the delivered HTML at all and 1 matched an element carrying no text.

21 min read Lantad

This run counted how many sites use it. Lantad requested the robots.txt of all 1,419 hostnames in this repository's two committed corpus seed files on 30 September 2026, then the home page of each, sent as its own declared crawler user agent with redirects followed and a twenty second timeout, from one network location, with no JavaScript executed. 1,126 answered HTTP 200 with an HTML content type. 14 of the 1,419 robots.txt files disallowed this scanner at the site root and 12 of those were among the 1,126, so they were dropped unread, leaving 1,114 pages. Because the property is defined for article pages and not only for front doors, a second pass took the 86 hostnames in the news and local media strata that answered and had internal links, followed up to three candidate article URLs on each, and kept the first page whose own markup declared an Article type. That identified an article page on 69 of them.

In short

  • Speakable schema is declared by almost nobody: Lantad read 1,114 home pages and 69 news article pages drawn from 1,419 corpus hostnames on 30 September 2026 and found the property on 8 home pages and 5 article pages, 11 distinct hosts in all.
  • Of the 12 speakable CSS selectors that name a single element by id, class or tag name and can therefore be resolved exactly from the delivered bytes, 4 matched no element and 1 matched an element carrying no text, measured by Lantad on 30 September 2026.
  • The two governing documents disagree on the property name: schema.org spells it xpath and lists three kinds of content locator, Google's page carrying Last updated 2026-09-08 UTC spells it xPath and lists two, and 4 of the 5 hosts using that form spelled it schema.org's way.
  • All 5 hosts using the XPath form on 30 September 2026 included the path to the head title element, which is what Google's own code example marks, while the guidelines on the same page ask for roughly two to three sentences of the story instead.
  • texastribune.org pointed the property at an element holding 3,667 words including a newsletter solicitation, and techcrunch.com pointed it at a purpose built element holding 59, both measured by Lantad on 30 September 2026.
StageCountWhat happened
Hostnames requested1,419The committed corpus, an editorial frame rather than a random draw
Answered 200 with an HTML content type1,126Of the rest, 212 answered 403, 33 answered 503 and 17 returned no status
Disallowed this scanner in robots.txt12Dropped without reading the page
Home pages read and analysed1,114The denominator for the home page figures
Carried parseable application/ld+json623A further 7 served a block that did not parse as JSON
Declared a WebPage or an Article family type206The two types schema.org says the property is used on
Declared the speakable property80.72 percent of the 1,114
News hostnames with an Article page identified69From a second pass over the news and local media strata
Of those, declared the speakable property57.2 percent of the 69
One GET of https://<host>/robots.txt and one of https://<host>/ for each of the 1,419 hostnames in worker/seeds/corpus-seeds-platform.json and worker/seeds/corpus-seeds-industry.json, sent as LantadBot/1.0 (+https://lantad.co/bot) with redirects followed and a twenty second timeout, from one network location, with no JavaScript executed, falling back once to the www subdomain where the apex did not answer with HTML. Measured by Lantad on 30 September 2026.

What is speakable schema and what actually reads it?

The property is specified in two places and they do not say quite the same thing, which turns out to matter later. The schema.org definition describes it as indicating sections of a web page that are particularly speakable in the sense of being highlighted as being especially appropriate for text to speech conversion, lists three kinds of content locator value, and states that it is used on the Article and WebPage types. The same page carries a usage band of 100 thousand to 1 million domains, attributed to monthly aggregations from Google's web index and dated August 2026. That band sounds large until it is placed against the others on the same scale: an earlier run reading those bands found that only 16 schema.org types reach 10 million domains, so a property sitting one or two bands below that is long tail rather than widely deployed, and the band is compatible with the rate measured here rather than in tension with it.

The consumer side is narrower than most people implementing it appear to believe. Google's documentation for the feature, carrying Last updated 2026-09-08 UTC, is titled for beta status and opens by saying the feature is in beta and subject to change. It states that the Google Assistant uses the markup to answer topical news queries on smart speaker devices, returning up to three articles from around the web with audio playback using text to speech, and that the property works for users in the United States that have Google Home devices set to English, and publishers that publish content in English. It adds that Google hopes to launch in other countries and languages as soon as a sufficient number of publishers have implemented it, which is a dependency running in the opposite direction from the one a site owner assumes.

So the honest scope is a smart speaker feature for English language news in one country, not a general instruction to answer engines. Nothing on either page says that any AI crawler reads the property, and this run observed no crawler doing anything: it made one request per URL and held no server logs. Lantad's own engine does not award a point for it either, and that is worth stating plainly rather than leaving to inference, because the grader here does score schema. The string speakable does not appear anywhere in core/src or worker/src, so a page that ships it perfectly gains nothing in the scoring methodology and a page that omits it loses nothing.

  • The property marks sections suited to text to speech Stated by the schema.org definition, which also names Article and WebPage as the types it is used on.
  • Google Assistant uses it for topical news queries Stated by Google's feature page, Last updated 2026-09-08 UTC: up to three articles on smart speaker devices with text to speech playback.
  • Availability is one country and one language Same page: users in the United States with Google Home devices set to English, and publishers publishing in English.
  • The feature is stable The page is titled for beta and states that the feature is in beta and subject to change, so requirements may still move.
  • An AI answer engine reads the property Neither page says so, and no vendor documentation read in this run addresses it. Undocumented rather than disproved.
  • Lantad scores it It does not. The string does not occur in core/src or worker/src, so it neither earns nor costs a grade.
What the two governing documents commit to about the speakable property, as read at source on 30 September 2026. Present means a published page states the behaviour, not that the behaviour carries weight anywhere.

8 of 1,114 home pages, and 5 of 69 news article pages

On the home pages the count is 8. Those eight sit in four strata and they are not the group anyone would predict for a news feature: etobicokerehab.com on a Wix or Squarespace build, agentplace.io among the single page application startups, spearmintlove.com in the Shopify direct to consumer stratum, auth0.com, netlify.com and tiendanube.com in the software stratum, and aljazeera.com and foxnews.com in the news stratum. Six of the eight are not news publishers at all, which is the first sign that the property is being deployed as a general purpose signal rather than as the smart speaker feature its documentation describes.

Three denominators are available and each says something different. Against all 1,114 pages read, 8 is 0.72 percent. Against the 623 that carried a parseable JSON-LD block, it is 1.3 percent, and that is the fairer comparison because a site with no markup at all had no opportunity. Against the 206 pages that declared a WebPage or an Article family node, the two types the vocabulary says the property is used on, it is 3.9 percent. The 623 figure is consistent with the wider picture from earlier runs, where 141 of 382 home pages carried no structured data in the raw HTML, and with the wider pattern in which a handful of schema.org types carry almost all deployment while the long tail of properties like this one sits far below.

The article pass matters more, because Google's page frames the feature around news content. 88 of the 160 news and local media hostnames were readable and 2 of those offered no usable internal link, leaving 86 to follow. 113 candidate pages returned HTTP 200 across them, and 69 hostnames yielded a page whose own markup declared an Article type. 5 of those 69 declared the property: aljazeera.com, cnbc.com, foxnews.com, techcrunch.com and texastribune.org. That is 7.2 percent, ten times the home page rate, so the feature is concentrated where it is meant to be even though the absolute number is small. One further host, techcrunch.com, carried an element with the id speakable-summary on its home page while declaring no speakable property there, so the target element ships site wide and the pointer to it appears only on articles.

The access picture is part of the finding rather than a footnote. The news stratum is the most heavily defended in the corpus: 62 of its 128 hostnames answered this identified crawler with HTTP 403, which is why 128 seeded hosts became 58 readable home pages. An earlier run measured the other half of that same wall, finding that 77 of 706 robots.txt files blocked a citation crawler and 42 of those were news sites. A feature whose whole premise is that publishers implement it, on a set of publishers who mostly refuse identified crawlers, has a distribution problem before any markup is written.

PopulationDenominatorDeclared speakableShare
All home pages read1,11480.72 percent
Home pages with parseable JSON-LD62381.3 percent
Home pages declaring WebPage or an Article type20683.9 percent
News hostnames with an Article page identified6957.2 percent
Distinct hostnames declaring it anywhere1,419 seeded110.78 percent
The speakable property by pass and denominator. Home page figures are of the 1,114 pages read; article figures are of the 69 news and local media hostnames where a page declaring an Article type was identified. Measured by Lantad on 30 September 2026.

5 of the 12 resolvable selectors pointed at no readable text

A declaration is a pointer, so the interesting question is not whether the property is present but whether the thing it points at exists. The 13 declarations found across both passes carried 31 content locator values between them, 19 of them CSS selectors and 12 in the XPath form. Every one of the 13 was re-fetched on a second independent request the same day and every figure below is read from that second capture, which matters because cnbc.com answered HTTP 200 during the first pass and HTTP 403 to a later hand check, so a single read would not have been enough to stand behind.

Of the 19 CSS selectors, 12 name a single element by id, class token or tag name and can therefore be resolved exactly against the delivered bytes. 4 of those 12 matched no element on the page at all. etobicokerehab.com points at the id block-70f96de2c69cdc3bdcbe, which appears in the document only inside the JSON-LD block that names it. agentplace.io points at a class named definition that no element carries. spearmintlove.com points at two classes, hero__heading and hero__subheading, neither of which is in the markup, and the first value in its list is the bare h1, which matched the site header heading and resolved to zero words of text because that heading wraps a logo link rather than a sentence. That is 5 of the 12 pointing at nothing a text to speech engine could read.

8 of the 12 matched an element and 7 of those held readable text, and those 7 divide into two failure modes and one success. netlify.com points at main, which resolved to 956 words, the whole content landmark of a marketing home page. texastribune.org points at the class entry-content, which resolved to 3,667 words, the full article body with a newsletter solicitation at the top of it beginning Never miss a story. Both sit directly against the guidance on Google's own page, which says that rather than highlighting an entire article with the markup a publisher should focus on key points, and recommends around 20 to 30 seconds of content per section, or roughly two to three sentences. The same site's entry-title selector matched 15 separate elements on the page, so it names the article's own headline and fourteen others without distinguishing them. techcrunch.com is the one implementation in this set that matches the guidance: a purpose built element with the id speakable-summary holding 59 words, which is about three sentences.

What this method cannot say is worth as much as what it can. No stylesheet was fetched and no JavaScript ran, so a class injected after load is recorded as absent and the four unmatched selectors are a floor rather than a certainty. 7 of the 19 CSS selectors were compound or attribute selectors and were not evaluated at all, among them the descendant forms on foxnews.com, cnbc.com and texastribune.org and two itemprop attribute selectors. All 12 XPath values were left unevaluated, because resolving them properly needs an XML view of the document this run did not build. The same caution applies here as to any claim about markup that names a value not present on the page: a selector that resolves in a browser may still fail in an extractor that never runs one, which is the gap the crawler view tool exists to show.

HostPageSelectorElements matchedWords of text
etobicokerehab.comHome#block-70f96de2c69cdc3bdcbe0No element
agentplace.ioHome.definition0No element
spearmintlove.comHome.hero__heading0No element
spearmintlove.comHome.hero__subheading0No element
spearmintlove.comHomeh110
agentplace.ioHomeh116
netlify.comHomeh116
netlify.comHomemain1956
aljazeera.comArticle.article-header133
techcrunch.comArticle#speakable-summary159
texastribune.orgArticle.entry-title1517
texastribune.orgArticle.entry-content13,667
Every CSS selector in the 13 declarations that names a single element by id, class token or tag name, resolved against the HTML of a second independent request. Word counts are of the text inside the matched element with script and style subtrees removed. Compound, attribute and XPath locators are excluded because this method does not evaluate them. Measured by Lantad on 30 September 2026.

schema.org spells the property xpath and Google spells it xPath

The two specifications diverge in three ways, and each divergence showed up in the measured markup. On the number of locator kinds, the schema.org page lists three: an identifier value URL reference, a CSS selector, and an XPath. Google's page lists two, CSS selectors and XPaths, and does not mention the identifier reference form at all. No site in this set used the identifier reference form, so the extra option in the vocabulary is unused here rather than contested.

On the name of the property, the two documents disagree outright. The schema.org page says to use the xpath property, all lower case. Google's page names the property xPath with a capital P, in its prose and in its code example. Of the 5 hosts using that form, 4 spelled it the schema.org way: aljazeera.com, cnbc.com, foxnews.com and tiendanube.com all ship xpath. Only auth0.com ships xPath. A consumer doing an exact key match against one document therefore misses four fifths of the real deployments or one fifth of them, depending which document it was written from, and neither page acknowledges the other spelling. This is the same class of defect as the syntax problems an earlier run found when 117 of 612 home pages with JSON-LD carried a validator defect, except that here the markup is correct against one published specification and wrong against the other.

On using both locator kinds at once, Google's page is explicit and says twice, once under each property, to use either cssSelector or xPath and not to use both. cnbc.com does exactly that. Its article carried a single SpeakableSpecification object holding an xpath array of two values and a cssSelector array of one, read from a second independent request to confirm it. That is one of the two clearest rule breaks in the set. The other is placement: agentplace.io attaches the property to an Organization node, and Organization is not among the types either document says the property is used on. aljazeera.com attaches it to a CollectionPage and foxnews.com to a node typed as both WebPage and CollectionPage, which are WebPage subtypes and so sit inside the stated domain.

None of this is exotic. It is the ordinary result of a property that has two governing documents, a beta label, and no validator anybody runs. Compare the position with a type that has a much larger deployed base and much clearer consumer behaviour: Google removed the FAQ rich result and Lantad still scores FAQPage, because the markup remains the most extractable question and answer shape there is. Speakable has the opposite profile. The consumer is named and narrow, the deployed base is tiny, and the two specifications do not agree on what to type.

Questionschema.org saysGoogle saysWhat the markup did
Kinds of content locatorThree: identifier reference, CSS selector, XPathTwo: CSS selectors and XPaths19 CSS selectors, 12 XPath, 0 identifier references
Name of the XPath propertyxpath, all lower casexPath, capital P4 hosts lower case, 1 capital
Using both kinds togetherNot addressedUse either one, do not use bothcnbc.com used both in one object
Types the property is used onArticle and WebPageThe Article or Webpage objectagentplace.io used Organization
How much text to markNot addressedTwo to three sentences, 20 to 30 seconds956 and 3,667 words on two hosts
Where the schema.org definition and Google's feature page disagree, and what the 13 measured declarations did. Both pages read at source on 30 September 2026; Google's page carries Last updated 2026-09-08 UTC. Measured by Lantad on 30 September 2026.

All five XPath users marked the head title, which is what the example marks

Google's page carries one code example for the XPath form, and the two values in it are the path to the head title element and the path to the content attribute of the meta description. Every one of the 5 hosts using that form on 30 September 2026 included the first of those two paths. 4 of the 5 also included the meta description path: aljazeera.com and cnbc.com with double quotes around the attribute value, tiendanube.com and auth0.com with single quotes exactly as printed in the example. auth0.com ships the example's two values and nothing else. tiendanube.com ships those two plus a path to any h1 and a path to the first paragraph. foxnews.com ships the title path plus a path to an h1 in the body.

The pattern is clear enough to name: these are not independent implementations that happened to converge, they are the documentation example in production. That would be unremarkable except for what the example marks. The title element and the meta description are not sentences of the story. They are the page's metadata, the same two fields a search result has always been built from, and an earlier run found that 38 of 1,083 home pages sent a crawler no name at all while most of the rest carry both fields without any thought about how they sound read aloud.

Two paragraphs above the code, the same guidelines ask for something different. They say content indicated by the markup must have concise headlines or summaries that provide comprehensible and useful information, that a publisher including the top of the story should rewrite it to break information into individual sentences so that it reads more clearly for text to speech, and that the recommendation is around 20 to 30 seconds per section or roughly two to three sentences. It also lists what to keep out: datelines giving the location where the story was reported, photo captions, and source attributions. A meta description written for a search snippet satisfies none of that by design, and a title element ending in a pipe and the publication name satisfies it less. foxnews.com marked a head title on its home page that reads as a list of section names separated by pipes, which is a reasonable search result and an unreasonable sentence.

This is why the techcrunch.com case is the useful one to copy. Marking an element built for the purpose, holding 59 words, is the only implementation here that does what the guidance describes rather than what the example shows. It is also the only one that would survive a redesign, because it depends on an id placed deliberately rather than on a class name from a theme. A selector tied to a theme class is a hostage to the next template change, which is the same fragility that makes 27 of 165 article pages declare a headline their own h1 does not contain.

How the documentation example reaches production, traced from what the 5 XPath users on 30 September 2026 actually shipped. A diagram of the observed pattern, not a claim about any publisher's intent.

What to check before shipping speakable markup

The measurement above suggests four checks, and all four can be run by hand on any page in a couple of minutes without a tool. The first is whether the element exists in the HTML the server sends, not in the page a browser builds. 4 of the 12 resolvable selectors in this set failed that check, and because this run executed no JavaScript it cannot say how many of the four a rendered inspector would have shown as present. Fetching the raw document and searching it for the id or class token is the test that matches what an extractor sees, and it is the same discipline behind reporting that 27 of 404 home pages gained structured data only when a browser ran the page.

The second is how much text the selector covers. 20 to 30 seconds is the published recommendation, and 3,667 words is not that. Counting the words inside the target element takes one command and it is the difference between marking a summary and marking a page. The third is which node carries the property. Article and WebPage are the stated domain, so a property hung on an Organization node is a statement the vocabulary does not define, in the same way that a missing language declaration was a statement 429 of 615 pages declined to make when none of their parseable JSON-LD named a language. The fourth is which spelling to ship, and here the honest answer is that there is no safe single answer: the deployed base favours lower case xpath four to one, the consumer whose behaviour is documented spells it xPath, and a publisher who cares can reasonably use the CSS selector form instead and avoid the question.

What none of these checks can establish is whether any of it pays. This post is a count of markup and the documented behaviour of one smart speaker feature, and it is not evidence that the property moves a citation in ChatGPT, Claude, Perplexity or Google AI Overviews. No engine other than Google Assistant is documented as reading it, no crawler was observed reading it here, and the corpus is an editorial sampling frame assembled for platform and industry coverage rather than a random draw, so every rate above describes these 1,419 hostnames and nothing wider. Anyone deciding where to spend an hour of markup work has better grounded options: the answer engine and AEO fundamentals that decide whether a page is readable at all, the entity signals that decide whether a brand is recognised, and the ordinary question and answer shapes that an engine assembling a response can actually lift, which is why 298 of 1,277 FAQ answers not being on the page was a more consequential finding than this one.

The reason to publish a null result like this is that the alternative is worse. A vocabulary property that sounds perfectly designed for answer engines, has a real specification, and shows a usage band in the hundreds of thousands of domains is exactly the kind of thing a checklist acquires and never loses. It is in the corpus at 0.72 percent of home pages, it is half broken where it appears, its two specifications disagree on its name, and its one documented consumer is a smart speaker in one country. Those are all things worth knowing before it goes on a roadmap, and they sit alongside the rest of the measured research rather than above it.

  • Element exists in the delivered HTML 4 of 12 failed Four selectors matched no element at all, and one matched an element holding no text. Search the raw document, not a rendered inspector.
  • Marked text is two to three sentences 2 of 7 overshot badly 956 words on netlify.com and 3,667 on texastribune.org against a published recommendation of 20 to 30 seconds per section.
  • Property sits on Article or WebPage 1 of 13 misplaced agentplace.io attached it to an Organization node, which is outside the domain both documents state.
  • Spelling matches the intended consumer No safe answer 4 hosts ship xpath, 1 ships xPath, and the two governing documents each print only one of them.
The four checks the measured defects suggest, with what this run found for each across the 13 declarations read on 30 September 2026.

Written by

Lantad

Published .

There is one property in the schema.org vocabulary whose entire job is to tell a machine which sentences to read out loud. It is called speakable, and speakable schema is the closest thing the markup vocabulary has to an instruction addressed at an answer engine rather than a search result: not what this page is about, not who published it, but read these words and skip the rest. If any part of structured data were going to matter to a system that answers a question in a sentence, this would be the candidate.

Common questions

What is speakable schema?

It is a schema.org property that marks the sections of a page best suited to being read aloud by text to speech. It takes a SpeakableSpecification value holding either a cssSelector or an XPath list that points at the elements to read, and schema.org states it is used on the Article and WebPage types.

Does speakable markup help me get cited by ChatGPT or Perplexity?

Nothing published says so. The only documented consumer is the Google Assistant, which Google's feature page describes as using the markup to answer topical news queries on smart speaker devices for users in the United States with Google Home devices set to English. No AI answer engine vendor documentation read in this run addresses the property, and this run observed no crawler reading it.

How many sites actually use speakable?

Very few. Lantad found it on 8 of 1,114 home pages and on 5 of the 69 news hostnames where an Article typed page could be identified, measured on 30 September 2026 across a 1,419 hostname corpus. The schema.org page reports a usage band of 100 thousand to 1 million domains from Google's web index in August 2026, which is one or two bands below the types that reach 10 million domains and is consistent with the rate measured here.

Should I write cssSelector or xPath?

The two specifications disagree, so check which consumer you are writing for. schema.org names the property xpath in lower case and Google's page names it xPath with a capital P. Of the five measured hosts using that form, four shipped lower case and one shipped Google's spelling. Using the cssSelector form instead avoids the question, and Google's page says to use one form or the other and not both.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.