BlogFindings

sameAs schema: 85 of 317 organizations named a reference entry, and 310 named a social profile

Lantad read the home page of 1,419 hostnames on 25 September 2026 with no JavaScript executed. 1,068 answered HTTP 200 with HTML, 433 carried an identity node the scanner recognises, and 317 of those declared sameAs on it. Those 317 pointed at 185 distinct hosts, of which the six largest are social networks: 310 named a social profile and 85 named a Wikipedia or Wikidata entry.

19 min read Lantad

This run counted. Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt, evaluated the site root for its own crawler token before requesting anything else, then requested the home page of every host that allowed it and parsed the raw bytes with no JavaScript executed. 1,068 answered HTTP 200 with an HTML content type. The question is not how many sites use sameAs. It is what the ones that do are pointing a machine at, and whether those targets identify anything a resolver outside the page could use. The answer, for a non-rendering AI crawler, is mostly Facebook.

In short

  • Lantad read 1,068 home pages on 25 September 2026 and found 433 carrying an identity node, of which 317 declared sameAs on that node and 116 did not, with 31 of the 116 carrying sameAs somewhere else on the page instead.
  • The sameAs schema property that those 317 sites actually shipped points overwhelmingly at social networks: facebook.com on 230 sites, linkedin.com on 220, youtube.com on 205 and instagram.com on 192, against en.wikipedia.org on 62 and wikidata.org on 35.
  • 85 of the 317 named a Wikipedia or Wikidata entry, 73 Wikipedia and 35 Wikidata with 23 naming both, and 227 named at least one social profile and no reference entry at all.
  • Of the 216 distinct Wikipedia and Wikidata URLs this corpus declared, 215 returned HTTP 200 when requested the same day and one, seota.com's Wikidata item Q135013201, returned 404.
  • 46 of the 2,437 sameAs values in the corpus were not absolute URLs, being 36 empty strings, 3 mailto addresses, 2 bare hostnames with no scheme and 5 others, and 7 sites still name a plus.google.com profile on a network Google shut down in April 2019.
StageCountWhat happened
Hostnames asked1,419The committed corpus, an editorial frame rather than a random draw
robots.txt disallowed LantadBot at the root48Never asked for a home page
Home pages requested1,37112 of the 48 were closed by a rule, 36 by a 5xx robots.txt
Did not answer 200 with HTML303216 answered 403, 34 failed before returning a status, 22 answered 429, 18 answered 503, 13 returned something else
Home pages read1,068The denominator for every figure below
Carried at least one JSON-LD block607461 carried none at all
Carried an identity node433Named, in the Organization family, and attributed to the site itself
Declared sameAs on that identity node317116 identity nodes declared none
One GET of https://<host>/robots.txt, then one GET of https://<host>/, each as LantadBot/1.0 with redirects followed, a fifteen second timeout on robots.txt and twenty on the home page, no JavaScript executed and from one network location. Measured by Lantad on 25 September 2026 across the 1,419 hostnames in worker/seeds/corpus-seeds-industry.json and worker/seeds/corpus-seeds-platform.json.

What is sameAs schema, and what do its two specifications say it is for?

Two published definitions govern this property and they do not describe the same thing. The schema.org definition of sameAs reads: "URL of a reference Web page that unambiguously indicates the item's identity. E.g. the URL of the item's Wikipedia page, Wikidata entry, or official website." That is a disambiguation instruction. The word doing the work is "unambiguously", and the three examples given are two encyclopedias and your own domain. Nothing in it mentions social media.

Google's Organization structured data documentation, carrying Last updated 2026-09-08 UTC, lists sameAs as a Recommended property and defines it as "The URL of a page on another website with additional information about your organization, if applicable. For example, a URL to your organization's profile page on a social media or review site. You can provide multiple sameAs URLs." That is a different instruction. It is "more about you", and its worked example is the social profile that schema.org's definition never raises.

Neither is wrong. They are the vocabulary's author and its largest consumer describing one property for two purposes, and a publisher reading either one in isolation comes away with a different idea of what to write. That divergence is the reason this measurement is worth taking: whichever definition a site followed is visible in the bytes, because a Wikidata item and an Instagram handle are not the same kind of claim. One says which entity this is among all the entities with similar names. The other says where else the entity posts.

Lantad treats the property the schema.org way. Its entity confidence diagnostic weights sameAs at 25 points of 100, which is the largest non-gate weight in that checklist, and the reason recorded in core/src/config.ts is exactly the definitional one: it is the only property in the set whose schema.org definition is an identity-resolution statement. That weight is a decision somebody made, not a measured effect. Nothing published by any answer engine says what a sameAs URL is worth, and this post does not claim otherwise. What follows is a count of what 1,068 servers returned.

schema.org, the vocabulary

  • URL of a reference Web page that
  • unambiguously indicates the item's
  • identity. E.g. the URL of the item's
  • Wikipedia page, Wikidata entry, or
  • official website.
  • Examples given: 2 encyclopedias, 1 own site

Google, the consumer

  • The URL of a page on another website
  • with additional information about your
  • organization, if applicable. For example,
  • a URL to your organization's profile page
  • on a social media or review site.
  • Example given: a social or review profile
The two published definitions of sameAs, quoted exactly as read at source on 25 September 2026. The schema.org page carries no dated revision marker; the Google page carries Last updated 2026-09-08 UTC.

317 of 433 identity nodes declared sameAs, and 461 pages carried no JSON-LD at all

The funnel matters more than the final ratio, because most of the loss happens before sameAs is even reachable. Of the 1,068 home pages read, 607 carried at least one JSON-LD block and 461 carried none, which is close to the 141 of 382 that an earlier and smaller run of this corpus found when it counted home pages carrying no structured data in the raw HTML. Carrying a block is not the same as declaring an identity: 433 of the 1,068 carried a node that the scanner will accept as the site's own identity, meaning a node in the Organization family that has a name and that is attributed to this site rather than to something the page merely mentions. So 174 pages published structured data that identified no publisher, which is the same shape as the finding that 103 of 385 pages with JSON-LD named no organization.

Among the 433 that did name themselves, 317 put sameAs on that node and 116 did not. The 116 are not all silent. 31 of them carried sameAs somewhere else in the same document, on a node that identifies an author, a product, a brand or a place rather than the site, which means the markup contains identifiers that attach to the wrong subject. ca.gov, hel.fi and maroc.ma are three of the 116 with nothing anywhere.

The 317 nodes carried 1,601 distinct absolute sameAs URLs between them. The median site declared five, the ninetieth percentile was nine, the largest single identity node carried seventeen, and 18 sites declared exactly one. That last group is worth listing because a single sameAs is the clearest possible statement of what a site thinks the property is for. dr.dk's one link is its Danish Wikipedia article. tec.mx's one link points back at tec.mx. The other sixteen are a social profile each: LinkedIn for cypher.build, carecycle.ai, cardamon.ai, impactorigins.co and trially.ai, Facebook for four veterinary, dental, legal and heating businesses, Instagram for yunastories.com, X for rappler.com, GitHub for countrystatecity.in and a Y Combinator company page for dollyglot.com.

Reading the other four signals in the same diagnostic against the same 1,068 pages puts sameAs in context. 358 pages declared a logo, 317 declared sameAs, 219 declared an absolute @id, and only 169 declared a machine-readable category for what the organisation is. The identity that most of this corpus publishes is a name and a picture.

  • At least one JSON-LD block 607 of 1,068 461 pages carried none
  • An identity node (named, Organization family) 433 of 1,068 174 pages carried JSON-LD that named no publisher
  • logo on the identity node 358 of 1,068 The most common non-gate property in this corpus
  • sameAs on the identity node 317 of 1,068 116 identity nodes declared none
  • An absolute @id on the identity node 219 of 1,068 What links several nodes into one graph rather than islands
  • A machine-readable category 169 of 1,068 What the organisation is, as a value rather than a sentence
Presence of each entity signal across the 1,068 home pages that answered HTTP 200 with an HTML content type, read as raw bytes with no JavaScript executed and assessed with assessEntityConfidence from core/src/entity.ts. Measured by Lantad on 25 September 2026. Presence is a count of what the markup declares, not a measure of what any engine does with it.

85 of 317 named an encyclopedia entry, and one of those entries no longer exists

The two targets schema.org names by example are Wikipedia and Wikidata, and 85 of the 317 sites named at least one of them. 73 named a Wikipedia article, 35 named a Wikidata item, and 23 named both. 227 sites named a social profile and no reference entry at all, which is the single clearest result in this run: on the reading published by the vocabulary itself, roughly seven in ten of the sites using sameAs are using it for something the vocabulary does not describe.

Every Wikipedia and Wikidata URL the corpus declared anywhere in its JSON-LD was then requested, 216 distinct addresses in all, as LantadBot with redirects followed. 215 returned HTTP 200. One returned 404: seota.com declares https://www.wikidata.org/wiki/Q135013201, and that item is not there. A sameAs pointing at a deleted Wikidata item is worse than no sameAs, because it is a specific, machine-checkable assertion that fails the check. It is also a good illustration of why this property rots differently from the rest of a page: the claim lives on your server and the evidence lives on somebody else's.

The distribution across the corpus is where the honest caveat sits, and it cuts against the simple reading. SaaS is the strongest stratum by a distance, with 33 of its 72 sameAs-declaring sites naming a reference entry, followed by news at 15 of 32 and finance at 12 of 28. Seven of the eighteen corpus categories produced none at all: the Framer, Shopify direct-to-consumer, Wix and Squarespace, Bubble, local media, static documentation and SaaS marketing strata declared 71 sites with sameAs between them and not one reference entry.

That is not negligence. A dental practice in Asheville has no Wikipedia article and no Wikidata item to link, so a Facebook page is not the wrong answer for that site, it is the only answer available. The finding is narrower and it holds anyway: for the organisations that do have a reference entry, mostly the large ones, naming it is free and 227 sites did not. The same asymmetry showed up when this corpus counted opening hours on 172 healthcare and travel home pages and found 2, and when it counted authors on 111 article pages: the property that a large organisation could complete cheaply is the one left empty. If you want the practical version for one engine, the notes on getting cited in Google AI Overviews cover what is and is not controllable from the page.

Corpus categoryPages readIdentity nodesameAsReference entry
saas117817233
education10219145
healthcare9528206
finance95392812
government8816112
travel7922184
ecommerce6226225
news59433215
wix-squarespace541210
spa-startups4424171
webflow421391
bubble-nocode42970
wordpress-smb4030121
Home pages read, identity nodes found, sameAs declarations and reference entries named, by the corpus category each hostname is filed under in the two committed seed files. Categories reading fewer than 40 hosts are omitted. Measured by Lantad on 25 September 2026. A reference entry means a URL on wikipedia.org or wikidata.org.

46 values were not URLs, and seven sites still link a network shut down in 2019

Across the whole corpus, 588 nodes carried a sameAs property and they held 2,437 values between them. 46 of those values are not absolute http or https URLs, so they identify nothing. 36 are empty strings, produced by a template that emitted the key whether or not a value existed. Three are mailto addresses, on downtownnotarytoronto.com, gentrydentistry.com and yogaoffeast.com, which is a reasonable misreading of a property whose name suggests "other places to find us". Two, both on messly.com, are bare hostnames with no scheme: www.facebook.com/messlyuk and www.twitter.com/messly_uk. Three, on adonis.io, are site-relative paths: /blog, /resources and /careers. One, on statefarm.com, is a LinkedIn URL with a stray double quote character at the front, which is a string concatenation escaping into the output. And one, on mskcc.org, is the string msk_sdc_block_cancer_types_select, which is a template variable name that reached production.

None of that is catastrophic and all of it is silent. A JSON parser accepts every one of those values, so the markup validates and the property is present and the identifier does not exist. It is the same class of defect this corpus found when it counted 117 of 612 home pages with JSON-LD carrying a defect, and the same reason a present field is not a working field.

Staleness is the other half. Seven sites still declare a plus.google.com profile: bhf.org.uk, medlineplus.gov, forbes.com, nzz.ch, scroll.in, scmp.com and italotreno.com. Google's own announcement of 10 December 2018 states that it decided "to accelerate sunsetting consumer Google+, bringing it forward from August 2019 to April 2019". Those identifiers have pointed at nothing for more than seven years, on the sites of a national heart charity, a United States government health library and three large newspapers.

The Twitter rename is the same rot in progress. Across the whole corpus rather than the 317 alone, 163 sites name twitter.com, 129 name x.com and 7 name both, so more of this corpus still points at the old domain than the new one. twitter.com currently redirects, so those links resolve, and a redirect is a working link rather than a broken one. It is still a corpus in which the most common non-Facebook identity claim is filed under a brand that no longer exists, which is what a field nobody revisits looks like from the outside. The head of the document has the same problem and this blog measured it yesterday, finding 38 of 1,083 home pages that sent a crawler no name at all.

  • 36 empty strings Identifies nothing A template emitted the sameAs key with no value behind it. The property is present and the identifier is absent.
  • 3 mailto addresses Identifies nothing On downtownnotarytoronto.com, gentrydentistry.com and yogaoffeast.com. An address is a contact point, not a reference page.
  • 3 relative paths Identifies nothing /blog, /resources and /careers on adonis.io. A relative reference only resolves against the page that carried it.
  • 2 hostnames with no scheme Identifies nothing www.facebook.com/messlyuk and www.twitter.com/messly_uk on messly.com. Real profiles, unusable as URLs.
  • 2 template leaks Identifies nothing A stray quote character prefixing a LinkedIn URL on statefarm.com, and the variable name msk_sdc_block_cancer_types_select on mskcc.org.
  • 7 sites naming plus.google.com Resolves to nothing Consumer Google+ was shut down in April 2019 per Google's announcement of 10 December 2018.
  • 163 sites naming twitter.com Resolves by redirect Against 129 naming x.com and 7 naming both. The links work; the brand they name does not exist.
  • 215 of 216 reference URLs Resolved HTTP 200 Every Wikipedia and Wikidata URL in the corpus was requested the same day. One 404: seota.com's Q135013201.
The 46 sameAs values in this corpus that are not absolute http or https URLs, grouped by what was written instead, plus the stale-target counts. Values quoted exactly as returned. Measured by Lantad on 25 September 2026 across 2,437 sameAs values on 588 nodes.

What this measurement does not show

No AI crawler fetched anything in this run. Every request came from LantadBot, so nothing here is evidence about what GPTBot, ClaudeBot or PerplexityBot does with a sameAs URL, and no vendor behind the crawler tokens this scanner evaluates documents it. Google's Organization page documents that the property is read for Search, not that a given value changes any generative result. The 25 point weight in this scanner's own entity checklist is a design decision recorded in a config file with its reasoning attached, and it is not a measured effect. Anyone telling you that adding a Wikidata link produces a citation is making a claim nobody outside those engines can currently test. The honest version is narrower: the property exists to disambiguate, most sites are not using it that way, and the cost of using it that way is one line.

Nothing here tests whether a declared identity is true. sameAs is an assertion by the site about itself, and this run checked that 216 encyclopedia URLs resolve, not that the entities behind them are the sites that claimed them. A site could name any Wikipedia article and this measurement would count it. Nor was any social profile requested: the 1,870 values include profiles that may be deleted, renamed or never to have existed, and only the reference entries were verified.

Only the home page was read on each host. An organisation's identity markup usually lives on the home page, which is why this is the right page to read, but a site that publishes sameAs on an about page and not on the home page counts here as declaring none. Every figure is a statement about home pages, and every count of an absence is a ceiling on what the site publishes rather than a proof of what it does not. No JavaScript was executed either, which is the point rather than a shortcut, because the question is what a crawler reading delivered bytes is given. A page that injects its Organization node from a script would be counted here as having none, and that is a real gap in the other direction: this run did not fetch a rendered view, so it cannot say how many of the 461 pages with no JSON-LD have some after their bundle runs. You can check a single page against this scanner's raw-bytes view with what GPTBot sees, and the grading rules are set out in the methodology.

Each page was requested once, from one network location, on one date. Every figure is a claim about 25 September 2026 and nothing after it, and the named sites are a list of what those deployments served that day. The 303 hosts whose home page did not answer are not missing at random, since a server that refuses an unknown crawler is likelier to refuse other automation, so the 1,068 read here lean towards sites that are open to being read. The corpus itself is an editorial sampling frame assembled for platform and industry coverage rather than a random draw of the web, so every rate supports a statement about these 1,419 hostnames and nothing wider, a constraint set out at length in the crawlability study. Where the markup is generated by a platform rather than by the site, the per-stack notes cover the mechanics, including for Shopify stores, and the platform-specific citation notes for one engine are in the guide to getting cited by Perplexity.

  • 317 of 433 declared sameAs Measured Read from the raw bytes of one request per host, assessed with the scanner's own entity module.
  • 85 of 317 named a reference entry Measured A URL on wikipedia.org or wikidata.org, counted once per site across Organization-family nodes.
  • 215 of 216 reference URLs resolved Measured Each requested once as LantadBot the same day, redirects followed. Resolution, not verification of the entity.
  • The two published definitions differ Documented Quoted from schema.org and from Google's Organization page, Last updated 2026-09-08 UTC. Neither is a Lantad finding.
  • sameAs is worth 25 of 100 here A decision A constant in core/src/config.ts with its reasoning recorded. A weighting this company chose, not a measured effect.
  • Whether any sameAs moves a citation Not measured No engine publishes it, no crawler was run here, and nothing in this post claims it.
What each figure in this post is, and what it is not. Written against the same run, Lantad, 25 September 2026.

Written by

Lantad

Published .

An answer engine that names your company has to decide which company it is. The markup property built for that decision is sameAs, and it is the one property in the schema.org vocabulary whose own definition is an identity statement rather than a description. It has been in the vocabulary for years, it costs a line of JSON, and almost nothing has been published about what sites actually put in it.

Common questions

What is sameAs schema used for?

schema.org defines sameAs as the URL of a reference web page that unambiguously indicates the item's identity, and gives Wikipedia, Wikidata and your own official website as its examples. Google's Organization structured data documentation, Last updated 2026-09-08 UTC, defines it more broadly as a page on another website with additional information about your organization, and gives a social media or review profile as its example. The two definitions are the reason sites use it so differently: in this corpus 310 of 317 sites followed the second reading and 85 followed the first.

Should sameAs point at Wikipedia and Wikidata or at social profiles?

If you have a Wikipedia article or a Wikidata item, naming it costs one line and it is the target the vocabulary itself gives as an example. Most sites do not have one: seven of the eighteen strata in this corpus produced no reference entry at all, because a small business has no encyclopedia entry to link. Social profiles are permitted by Google's definition and are what 310 of the 317 declaring sites use. Nothing published by any answer engine says what either kind of target is worth, so this is a question of which definition you are writing to, not of a measured payoff.

How many sites actually declare sameAs?

In this corpus, 317 of the 1,068 home pages read on 25 September 2026, which is 29.7 percent of all pages read and 73.2 percent of the 433 that carried an identity node at all. The bigger loss happens earlier: 461 of the 1,068 carried no JSON-LD whatsoever, and a further 174 carried JSON-LD that named no publisher. A site cannot declare sameAs on an identity it has not declared.

Does a broken sameAs URL matter?

It is a specific, machine-checkable claim that fails the check, which is a different condition from having no claim. This run found 46 of 2,437 values that are not absolute URLs at all, including 36 empty strings and a template variable name, and one Wikidata item, seota.com's Q135013201, that returns 404. Whether any engine penalises that is not published by anyone and was not measured here. What can be said is that the value costs nothing to check and nobody appears to be checking it.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.