BlogFindings

Entity SEO: 107 of 163 schema topic statements named the site itself

Lantad asked all 1,419 hostnames in this repository's committed corpus for robots.txt on 8 October 2026 and read each home page it was allowed to read, with no JavaScript executed. 1,074 answered HTTP 200 with HTML and 612 of those carried parseable JSON-LD. 129 of the 612 used about, mainEntity or knowsAbout to name a subject, and of the 163 about and mainEntity values across 104 sites, 107 pointed at the site's own organisation, service or founder rather than at a topic.

18 min read Lantad

That makes about the most direct instrument a page has for telling an AI crawler what it is for. So this run went looking for it. One request per hostname for robots.txt, one for the home page where the file allowed it, no JavaScript, every JSON-LD block parsed and flattened with this scanner's own parser. The question was narrow: how many sites answer the question at all, and when they answer it, what do they say?

The answer to the first half is 129 of 612. The answer to the second half is the finding. Across the 163 about and mainEntity values on the 104 sites that declared either, 107 pointed at the site's own organisation, service or founder. The property that exists to name a page's subject was overwhelmingly used to name the publisher.

In short

  • Entity SEO asks a page to name its subject in markup, and 129 of the 612 home pages that carried parseable JSON-LD on 8 October 2026 named one using about, mainEntity or knowsAbout.
  • 107 of the 163 about and mainEntity values measured on 8 October 2026 resolved to a node carrying the site's own name, so the property that states a page's subject was used to state the publisher's identity.
  • On 88 of the 104 sites declaring either property, every value named the site itself, and on 12 of the 104 no value did.
  • 115 of the 116 bare @id references resolved to a node in the same page graph and 113 of those nodes carried a name, so the pattern was wired correctly and still said the wrong thing.
  • Google's Organization documentation, carrying Last updated 2026-09-08 UTC, lists no required properties and names neither about nor knowsAbout among its recommended ones.
  • Hostnames in the committed corpus 1419 392 platform strata, 1,027 industry strata
  • Home pages that answered 200 with HTML 1074 12 closed the root to LantadBot, 215 answered 403, 53 answered 503, 35 failed in the client
  • Carried at least one parseable JSON-LD block 612 1,151 blocks holding 8,731 nodes
  • Named a subject with about, mainEntity or knowsAbout 129 185 carried at least one of the five properties measured
  • Named a subject that was not the site itself 12 of the 104 sites declaring about or mainEntity
From 1,419 hostnames to 129 that named a subject. Measured by Lantad on 8 October 2026, one request per hostname for robots.txt and one for the home page, redirects followed, no JavaScript executed.

What is entity SEO, and what does it ask your markup to say?

It asks for two separate statements and most sites only make the first. The first is identity: this site belongs to this organisation, which has this name, this logo and these profiles elsewhere. That one is well covered, and we have measured it from several directions: 412 of 1,080 home pages named their site in neither source Google reads first, and 85 of 317 organizations named a reference entry in their sameAs list.

The second statement is subject matter, and it is a different claim. An engine that has resolved who you are still does not know what any given page of yours is for. Four properties in the vocabulary carry that information, and they divide the work cleanly. About names the subject of a document. MainEntity, defined as the primary entity described in some page or other CreativeWork, names the one thing a page is principally a page about, and its inverse is mainEntityOfPage. KnowsAbout is the odd one out: it hangs off an Organization or a Person rather than a document, and the specification is careful to say it indicates a topic that is known about, suggesting possible expertise but not implying it. IsPartOf, defined as an item that this item is part of, places a document inside a larger one rather than describing it.

The distinction matters because only the first two are statements about the page in front of the crawler. The counts below keep them apart for that reason, and they also exclude one case that would otherwise inflate mainEntity badly. On a FAQPage or a QAPage, mainEntity does not hold a topic: it holds the questions. We found 51 such nodes on 50 sites and took them out of every mainEntity figure in this post, because counting them would mean counting the FAQ markup we have measured separately as a topic statement.

PropertyWhat schema.org defines it asDeclared onSitesValues or nodes
isPartOfAn item that this item is part ofCreativeWork138170
aboutThe subject matter of an objectCreativeWork and five others97122
knowsAboutA topic that is known about, suggesting possible expertiseOrganization, Person39491
mainEntityOfPageInverse of mainEntityThing37201
mainEntityThe primary entity described in some pageCreativeWork1841
The four properties, their schema.org definitions as published, and how many of the 612 home pages carrying parseable JSON-LD declared each. Measured by Lantad on 8 October 2026.

How many sites named a subject at all?

129 of the 612, which is one in five of the pages that had already gone to the trouble of publishing structured data. The denominator is worth holding on to. These are not sites that ignore markup. They had 1,151 parseable JSON-LD blocks between them holding 8,731 nodes, which is a mean of just over fourteen nodes a page. They are the engaged minority. We have measured the other side of that line before: 141 of 382 home pages carried no structured data at all in the raw HTML.

Within the engaged minority the four properties are not adopted evenly, and the ordering is the opposite of what their definitions would suggest. The property declared most often is isPartOf, on 138 sites, which says the least: it places a page inside a site without describing either. The property declared least often is mainEntity, on 18 sites, which says the most. About sits between them on 97 sites. A vocabulary where the most informative property is the least used is a vocabulary being filled in by something other than an author with an intention.

The strata make the same point from another angle. Among the 89 software-as-a-service hosts carrying JSON-LD, 23 named a subject. Among the 52 news hosts, 13 did. Among the 26 government hosts, one did. The pages least likely to say what they are about were the ones published by institutions whose whole output is subject matter, and the pages most likely to say it were commercial sites whose subject is usually themselves. That is a hint about where these properties come from, and the next section turns it into a count.

  • isPartOf 138 places a page inside a site, describes neither
  • about 97 122 values
  • knowsAbout 39 491 values, and three sites hold 178 of them
  • mainEntityOfPage 37 201 nodes
  • mainEntity 18 41 values, after removing 51 FAQ and QA nodes on 50 sites
Sites declaring each property, of the 612 home pages that carried parseable JSON-LD. FAQPage and QAPage mainEntity nodes are excluded from the mainEntity row. Measured by Lantad on 8 October 2026.

107 of 163 topic statements named the publisher, not a topic

This is the measurement the post exists for. Taking every about and every non-FAQ mainEntity value across the corpus gives 163 values on 104 sites. For each one we resolved what it actually points at, read the target node's type and name, and compared that name against the names the same page gave its own Organization, WebSite, LocalBusiness or Person nodes. 107 of the 163 matched the site's own name.

The type distribution says the same thing without needing the name comparison. 73 of the 163 targets were typed Organization outright. Another 29 were a more specific organisation or person type: Service, Corporation, NewsMediaOrganization, MedicalOrganization, AccountingService, HVACBusiness, ProfessionalService, Person, and in one case an Organization that was also a BankOrCreditUnion. 102 of the 163 values, in total, pointed at an organisation or a human being. Only 16 pointed at a Thing, which is the generic type you reach for when the subject is a subject rather than a party.

Read site by site rather than value by value it is starker. On 88 of the 104 sites, every single about or mainEntity value named the site itself. On 12 of the 104, none did. The middle ground, where a site names itself on one page element and a topic on another, holds four sites. This is not a population of authors making a mixed job of a hard property. It is two populations: one using about to repeat its own identity, and a much smaller one using it as specified.

The twelve are worth naming because they show what the property looks like when it works. Bettersheets.co points about at a SoftwareApplication named Google Sheets, which is a genuine statement that the site is about somebody else's product. Etobicokerehab.com points at three Thing nodes named Chiropractic, Chiropractic treatment techniques, and Chiropractic and Manual Therapies. Trustmrr.com points about at a Dataset and mainEntity at an ItemList named Top startups by verified revenue, which is an accurate description of what that page is. Bonhomiegin.com points at a Product named Gin Bonhomie. None of those four sites gains a rich result from the markup. They are describing their pages for something that reads descriptions.

Target @typeValuesWhat the type describes
Organization73the publisher as a party
Thing16a subject, generically typed
WebContent16a block of content on the page
SoftwareApplication10a product, sometimes another company's
WebPage7another page, all seven being navigation
Service6a service the publisher sells
Corporation4the publisher as a party
ItemList4an enumerated list on the page
Everything else2721 further types, none of them on more than three values
What the 163 about and mainEntity values resolved to, by the target node's declared @type. Measured by Lantad on 8 October 2026 across the 104 corpus home pages declaring either property.

The references resolve, which is what makes this worth saying

There is an obvious explanation for the figures above that happens to be wrong, and ruling it out took a second pass over every page. 116 of the 163 values are not inline objects at all. They are bare references of the form about with a single @id key and nothing else, which is the compact way to point at a node defined elsewhere in the same document. A bare reference that points at nothing is a dead end, and if most of them dangled then the finding would simply be that the markup is broken, which is a far less interesting claim and one we have already made in another form: 117 of 612 home pages with JSON-LD carried a defect.

So we refetched all 129 sites, rebuilt each page's whole graph across every JSON-LD block on it, collected every @id the page defines, and checked each reference against that set. 115 of the 116 resolved. 113 of those resolved to a node that carried a name. Exactly one reference dangled.

That result matters more than it looks. It means the pattern is not an accident of tooling, not a half-finished migration and not a validator error. The pages are internally consistent. Somebody, or something, wired about to a real node with a real name, and the node it chose was the publisher's own. The markup is doing exactly what it was told to do. We have written before about the mechanics these references depend on, since 313 of 615 home pages gave no schema node an @id and a reference needs one at the other end; this corpus is the half that got that part right.

Two smaller patterns in the same data are worth recording because they show the property being used for things that are not subjects at all. Seven of the WebPage targets belong to one design studio whose home page declares seven mainEntity values pointing at pages named Work, Expertise, About, Latest, Awards, Contact and Clients. That is the navigation menu, declared as the primary entity the page describes. All 16 of the WebContent targets belong to a single cancer centre whose mainEntity values name promotional blocks, among them Award-Winning Excellence, Year After Year and 3X Match: Breast Cancer Awareness Month. In both cases the property is being used as a container for page furniture rather than as a statement of subject.

How a bare @id reference resolves, traced on povio.com as captured on 8 October 2026. The reference is valid at every step and the entity it arrives at is the publisher.

knowsAbout is a list, and three sites hold 178 of the 491 entries

KnowsAbout behaves differently from the other three and deserves its own count, because it is the only one of the four that is explicitly about expertise rather than subject matter. It hangs off the Organization or Person node rather than off the document, and the specification's wording is unusually hedged: a topic that is known about, suggesting possible expertise but not implying it, and it adds that no skill levels are distinguished.

39 sites declared it, carrying 491 values between them. The distribution is the point. One static analysis vendor declares 140 values on its home page alone. The top three sites hold 178 of the 491. The median across the 39 is ten. Two sites declare exactly one. A property with a median of ten and a maximum of 140 is not being used to make a claim; it is being used to enumerate a taxonomy.

The shape of the values says the same. 302 of the 491 are plain strings, which the specification permits, since knowsAbout accepts Text as well as Thing and URL. Another 142 are bare URL strings. Only 47 are objects with a name, and only 24 of those carry a sameAs, which is the property that would let an engine resolve the topic to a known entity rather than guess at a phrase. So of 491 declared areas of expertise across the corpus, 24 are tied to anything an engine could look up. That ratio is the same one we found when we counted identity references: a property that supports external resolution is mostly used without it. Entity confidence is the thing these links are supposed to buy, and a free text topic string buys very little of it.

None of this makes knowsAbout a mistake. A list of a hundred and forty topics is a reasonable thing for a developer tools company to publish, and it is more information than the 483 sites that named no subject at all. It is just not the same act as naming what one page is about, and a count that pooled the two would hide the gap this post is measuring.

  • sonarsource.com 140 28 percent of every knowsAbout value in the corpus
  • seota.com 22
  • trueform.agency 16
  • codehooks.io 15
  • povio.com 13
  • roberthalltaxes.com 13
  • confluent.io 13
  • crayo.ai 12
  • thefurrow.tv 12
  • diabetes.org 12
knowsAbout values per site, the ten highest of the 39 corpus sites declaring the property. The 39 sites carry 491 values between them, median ten. Measured by Lantad on 8 October 2026.

Where these properties come from, and why that explains the pattern

157 sites declared isPartOf or mainEntityOfPage, the two properties that place a document rather than describe it. On 69 of those 157, the page also carried the string that a single widely used WordPress search engine optimisation plugin leaves in its output. Reading the generator meta element on the same pages adds the rest of the picture: 22 of them declared WordPress and the plugin together, 14 carried the plugin with no generator element at all, 14 ran Google Site Kit beside it, seven ran a Divi theme and five ran WPML.

That is the mechanism. These properties arrive as a side effect of installing software, not as an editorial decision, and the software emits the pattern that is cheap and safe for it to emit. A plugin running on a home page knows two facts for certain: the site has an organisation, and this document belongs to that site. It does not know what the page is about, because nothing on the page tells it in a machine readable way. So it writes isPartOf pointing at the WebSite node and about pointing at the Organization node, both of which are true, neither of which answers the question the property was defined to answer. This is the same mechanism we found when 27 of 165 article pages declared a headline the page's own h1 does not contain: a generator filling a required slot with the value it can reach.

It also explains the stratum ordering from earlier. Commercial sites run plugins and therefore declare these properties by default, with the publisher as the subject. Government sites run bespoke stacks and therefore declare nothing. Neither population is making a claim about subject matter, and only one of them looks like it is.

Google's own documentation is consistent with treating the pattern as harmless rather than useful. Its Organization structured data page, Last updated 2026-09-08 UTC, lists no required properties at all, and its recommended list runs from name, logo, sameAs and description through address to seven registration identifiers including duns, leiCode and vatID. Neither about nor knowsAbout appears anywhere on it. The one place Google requires mainEntity is the Q and A page documentation, Last updated 2026-09-08 UTC, which states that the Question for the page must be nested under the mainEntity property, which is the container use rather than the topic use. Of the 25 features in Google's structured data gallery, read on 8 October 2026 and carrying Last updated 2026-06-15 UTC, none turns about into a search feature. The properties are vocabulary rather than rich result inputs, which is precisely why they are interesting to anyone optimising for something that reads vocabulary.

What the page declaredSites
Plugin fingerprint, generator names WordPress22
Plugin fingerprint, no generator element14
Plugin fingerprint, alongside Google Site Kit14
Plugin fingerprint, alongside a Divi theme7
Plugin fingerprint, alongside WPML5
Plugin fingerprint, alongside something else7
No fingerprint, no generator element62
No fingerprint, generator names WordPress8
No fingerprint, another generator18
The 157 corpus sites declaring isPartOf or mainEntityOfPage, by what the page's own generator meta element and plugin fingerprint declared. Measured by Lantad on 8 October 2026.

What this measured, and what it did not

One page per hostname, which means a home page. A home page is the page for which the publisher genuinely is the subject, so the finding is measured on the most forgiving possible sample. An article page whose about pointed at its publisher would be a clearer error than anything counted here, and this run did not look at article pages. If the pattern is a plugin default then it will appear on interior pages too, but that is an inference from the generator counts rather than something measured, and we have not measured it.

No JavaScript was executed, so a site that injects schema from a script was recorded as carrying whatever its raw HTML carried. That is a deliberate choice and it is also a real gap: 27 of 404 home pages with no structured data in the HTML gained some when a browser ran the page, and the same could be true of these properties. Everything here is what a crawler that does not render sees, which is the population most AI crawlers belong to.

One network location, one request per URL, one date. Rate limits and bot defences remove sites unevenly, and the 215 home pages that answered 403 are not a random sample of the corpus: they skew towards sites behind the defences that also tend to run the most deliberate markup. The 612 denominator is the pages we could read, not the pages that exist.

The name comparison has a known weakness in both directions. Matching a target node's name against the page's own Organization and Person names will count a subsidiary or a parent as the site itself when the names coincide, and will miss a self reference written in a different form, such as a legal name against a trading name. We ran it case insensitively on trimmed strings and nothing more clever than that. The type distribution is the check on it: 102 of 163 targets were an organisation or person type, which is an independent route to the same conclusion and does not depend on names matching at all.

Finally, nothing here measures effect. We counted what pages declare, not what any engine does with it, and no part of this run asked an engine anything. We have been explicit about that limit elsewhere and it applies with full force to a property no documented search feature consumes. Whether naming a real subject in about changes how often a page is cited is a question about AI visibility that this measurement cannot answer, and we would rather say so than imply a mechanism we have not tested. The methodology page states the general version of that rule, and the scanner's own bot page documents the agent that made these requests.

  • Declaration counts Measured 129 of 612 JSON-LD home pages named a subject with about, mainEntity or knowsAbout.
  • Reference integrity Measured 115 of 116 bare @id references resolved in the same page graph, 113 to a named node.
  • Self reference rate Measured 107 of 163 values named the site's own name, and 102 of 163 targets were an organisation or person type.
  • Interior pages Not measured One page per hostname. The generator counts imply the pattern repeats, and no interior page was read.
  • Script injected markup Not measured No JavaScript executed, so markup written by a script after load was not counted.
  • Effect on citation Not measured No engine was asked anything. Nothing here shows that naming a real subject changes an answer.
What this run established and what it left open. Measured by Lantad on 8 October 2026.

Written by

Lantad

Published .

Entity SEO is the practice of getting a machine to agree with you about what a page is, who published it and what it concerns. Structured data is where that agreement is written down, and the schema.org vocabulary contains a property whose entire job is to answer the last of those three questions. It is called about, and the specification defines it in six words: the subject matter of an object. Its expected value is a Thing, which in schema.org terms means an entity rather than a string, and its inverse is subjectOf.

Common questions

What is entity SEO?

It is the practice of making a page's identity and subject matter explicit enough that a machine resolves them the same way a reader would. In structured data terms it means naming the organisation or person behind the site, linking that entity to references an engine already holds, and separately naming what each page is about. The first half is well adopted. The second half is the gap this measurement found: 129 of 612 corpus home pages carrying parseable JSON-LD named a subject at all on 8 October 2026.

Is it wrong for a home page to point about at its own organisation?

On a home page it is defensible, because the publisher genuinely is what the page is about. The reason it is still worth counting is that the same markup is emitted by platform software on every page of a site, where it stops being true. 69 of the 157 corpus sites declaring isPartOf or mainEntityOfPage also carried a plugin fingerprint, which is the signature of a default rather than a decision.

Does Google use the about property?

Not as a documented search feature. Google's Organization structured data documentation, Last updated 2026-09-08 UTC, lists no required properties and does not name about or knowsAbout among its recommended ones. The only place its gallery requires mainEntity is the Q and A page, where the property must hold the page's Question rather than its topic. These properties are vocabulary a reader can consume, not inputs to a rich result.

How can I tell what my own pages declare?

Read the JSON-LD your server sends rather than what a browser assembles, flatten every block into a node list, and for each about or mainEntity value follow the @id to the node it names and look at that node's type and name. If it is your Organization node on a page that is not about your organisation, the property is filled in rather than answered. Doing it without rendering matters: this run executed no JavaScript, because that is what a non-rendering crawler sees.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.