BlogFindings

Person schema: 111 of 1,079 sites named a human, and 162 of those names appeared nowhere else on the page

Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt on 6 October 2026 as LantadBot, read every permitted home page with no JavaScript executed, then followed each page's own link to an about, team or company page and read that too. 1,079 home pages and 709 about pages answered HTTP 200 with HTML. 111 of the 1,079 sites declared a Person node on one of those two pages, 418 declared an organisation and no person at all, and of the 518 named Person nodes found, 272 carried neither of the two properties Google's own documentation names for working out who an author is.

19 min read Lantad

Person schema is that type. schema.org defines Person as a person, alive, dead, undead or fictional, and the vocabulary attaches jobTitle, worksFor, sameAs and knowsAbout to it, which between them say what someone does, who they do it for, and which other records on the web are about the same individual. Nothing requires a site to publish any of it. So Lantad went and counted how many do. All 1,419 hostnames in this repository's two committed corpus seed files were asked for robots.txt on 6 October 2026 as LantadBot, every permitted home page was read with no JavaScript executed, and then each home page's own link to an about, team or company page was followed and that page read the same way. 1,079 home pages and 709 about pages came back with HTTP 200 and an HTML content type. 111 sites named a human in markup on either of them.

In short

  • Person schema is the one construct that tells a machine a named human stands behind a page, and on 6 October 2026 Lantad found it on 90 of the 1,079 readable home pages in this repository's committed corpus and on 71 of the 709 about pages those home pages linked to, which is 111 sites in total.
  • 518 of those 1,079 home pages declared an Organization-family node and 418 of the 518 named no person anywhere across the two pages read, so the commonest shape in the corpus is a company that identifies itself to a machine and identifies nobody in it.
  • Of the 518 Person nodes that carried a name, 272 carried neither url nor sameAs, the two properties Google's article structured data documentation, last updated 2026-09-08 UTC, says it can use when disambiguating authors.
  • 162 of those 518 names appeared nowhere in the delivered HTML outside the JSON-LD block itself, so on 57 hostnames the structured data was the only place on the page where a crawler could learn the person existed.
  • Education declared a person on 1 of 102 sites, government on 1 of 86 and healthcare on 3 of 93, against 25 of 117 in the SaaS stratum, and the gap survives controlling for whether the site declared any organisation markup at all.
StageHostsWhat happened
Hostnames asked for robots.txt1,419392 in ten platform strata, 1,027 in eight industry strata
Never returned a status29Failed in the client, at DNS, or on the timeout
Returned a file that parsed into rules1,117220 more answered 4xx, which RFC 9309 reads as no restrictions
Permitted LantadBot at the site root1,32453 of the 66 refusals were a 5xx read as disallow, 13 were a rule
Home page answered 200 with HTML1,079The denominator for every rate below
Carried a JSON-LD block that parsed620625 carried a block; on 5 pages no block parsed
Declared an Organization-family node518Including Corporation, LocalBusiness and NewsMediaOrganization
Declared a Person node on the home page908.3 percent of the 1,079
Linked an about, team or company page718Same host, found from the home page's own anchors
That page answered 200 with HTML7099 answered a non-200 status
Declared a Person node on the about page7110.0 percent of the 709
Named a person on either page11110.3 percent of the 1,079 sites
Two GET requests per hostname at most, plus the robots.txt fetch, as LantadBot/1.0 (+https://lantad.co/bot), redirects followed, 20 to 25 second timeouts, no JavaScript executed, from one network location. Measured by Lantad on 6 October 2026 across the 1,419 hostnames in this repository's two committed corpus seed files.

What is Person schema for, and what does Google say it does?

A search engine or an answer engine handed a page about a company has to decide what kind of thing it is looking at and which entities the page is making claims about. The organisation half of that is well served in this corpus: 518 of the 1,079 home pages declared a node of an Organization-family type, and earlier runs over the same corpus have measured the logo those nodes declare and the telephone numbers they leave out. The human half has one construct and it is Person.

Google's documentation is specific about what it does with it. The article structured data reference, carrying Last updated 2026-09-08 UTC, lists author as a recommended property and states in its own words that there are no required properties. Its best practices section then says four things that this measurement can check. Make sure that all the authors that are presented as authors on the web page are also included in markup. When specifying multiple authors, list each author in their own author field. Use the Person type for people, and the Organization type for organizations, and do not use the Thing type. And in the author.name property, only specify the name of the author, with the publisher name, the job title and any honorific prefix moved to their own properties.

On the identifying properties it is equally direct: author.url is a link to a web page that uniquely identifies the author, and the page states that Google can understand both sameAs and url when disambiguating authors. That sentence is the reason this post counts those two properties rather than counting names. A name on its own is a string. Two people can share it, and an engine holding a string has no way to tie the byline on your page to any record it already keeps, which is the whole job of entity confidence as this scanner models it.

None of this is specific to Google. schema.org publishes the vocabulary, version 30.1 at the time of reading, and the same markup is what every engine reading the page gets. The point of measuring it is narrower than a ranking claim and worth stating plainly before any number: Person markup is a statement a page makes about a human, in a form a machine can act on, and a page that makes no such statement has left the engine to infer the human from prose. An engine that infers may infer nothing.

The two routes by which a machine reading one page can learn that a named human stands behind it, and where the corpus loses each one. Mechanism, not a measurement of any engine's behaviour.

How many sites declared Person schema at all?

90 of the 1,079 home pages carried at least one Person node. 71 of the 709 about pages did. Because they are not the same sites, the union is larger than either: 111 hostnames declared a person on one page or the other, which is 10.3 percent of the corpus that answered. The complement is the number worth sitting with. 518 home pages declared an organisation, and on 418 of those 518 no person was declared anywhere across the two pages read.

That shape is consistent with everything else this corpus has produced about identity markup. A previous run found 412 of 1,080 home pages that named no website in their own markup, and another found 313 of 615 pages that gave no node an identifier at all. Organisation markup is widely deployed and thinly populated, and the people are the thinnest part of it.

The counts are small enough that distribution matters more than the total. 524 Person nodes were found in all, 342 on home pages and 182 on about pages, and 142 of the 524 came from one hostname: eltiempo.com, a Colombian newspaper that marks up its reporters. The other 110 sites supplied 382 between them, the median site supplied 2, and 27 sites supplied exactly one. 357 distinct names appeared across the whole set. Any sentence about what the average site does here is a sentence about a handful of sites, which is why the rates above are reported against pages rather than against nodes.

Where a person did appear, one property on the organisation was usually the reason. 60 of the 1,079 home pages carried an Organization-family node declaring founder or founders, and 55 of those 60 also carried a Person node on the same page. Only 5 of the 1,079 home pages declared employee or member on an organisation. So the human who makes it into a company's structured data is the founder, and almost nobody else does. Whether that is a fair picture of the company is not a question markup can answer, but it is the picture an engine reading the structured data receives.

  • Home pages read 1079 pages Permitted by robots.txt and returning HTML
  • Home pages with an Organization-family node 518 pages 48.0 percent of the 1,079
  • Home pages with a Person node 90 pages 8.3 percent of the 1,079
  • About pages with a Person node 71 pages Of 709 about pages read
  • Sites naming a person on either page 111 pages 10.3 percent of the 1,079 sites
  • Organisation declared, no person anywhere 418 pages Of the 518 that declared an organisation
Pages and sites declaring each kind of identity node, out of the 1,079 corpus home pages that answered HTTP 200 with HTML and the 709 about pages reached from them. Measured by Lantad on 6 October 2026.

272 of 518 named people carried neither url nor sameAs

Every Person node found was read for the properties that do work beyond the name. 518 of the 524 carried a name and 6 carried none at all, three of them on getdianahr.com, whose about page declares one node typed Person with a name and a LinkedIn sameAs and three more carrying only a jobTitle and a description: HR Compliance Specialist, Payroll Specialist, Benefits Advisor. Those three are roles rather than people. The same site's home page declares a fourth node typed Person whose name is SMB Founder (Verified on G2), which is a review attribution rather than a human.

Of the 518 named nodes, 193 carried url, 70 carried sameAs, and 17 carried both, leaving 272 with neither. That is the finding this section exists for, because it maps directly onto the sentence in Google's documentation about disambiguating authors: on 272 of 518 declarations, a machine is handed a name and no address at which to resolve it. 157 carried jobTitle, 224 carried an image, 69 carried worksFor and 56 carried an @id that would let another node in the same graph point at them. 151 of the 524 nodes carried name as their only property, and 156 carried none of jobTitle, sameAs, worksFor, url or image.

The type errors are rarer than the thin descriptions and they are more specific. 11 of the 524 Person nodes carried a name identical to the name of an Organization-family node declared on the same page, which is the exact confusion Google's best practices warn against when they say to use Person for people and Organization for organizations. They come from five hostnames: redis.io six times, spectator.co.uk twice, and estesvalleyvoice.com, tucsonspotlight.org and busy.in once each, in every case a publication or a company typed as a person. Two more names arrived carrying an HTML character reference that nothing will decode, because the content of a script element is raw text: therightaccompany.com declares a Person named A & L Heating and Air, and dawn.com declares one named The Newspaper's Staff Reporter.

Eight names began with Dr., across five hostnames including northashevillefamilydentistry.com, smileisleorthodontics.com and predoc.ai, and that prefix is what the same documentation asks to be moved to honorificPrefix. That property was present on 5 nodes in the whole corpus. This is the least consequential defect on the page and it is worth one sentence because of what it implies about the rest: the honorific is the credential these practices sell, it is the part of the name a reader is meant to notice, and it was typed into the field reserved for the name because no plugin told anybody otherwise. Defect counting of this kind is what the schema markup validator run on this corpus was built to do, and the pattern repeats: the error is almost never a misunderstanding of the subject, it is a field doing a job it was not given.

PropertyNodesWhat it supplies to an engine
name518The string, and nothing resolvable on its own
image224A portrait, which identifies nobody to a parser
url193A page that uniquely identifies the person
jobTitle157What they do, unresolved against any organisation
sameAs70Another record on the web about the same person
worksFor69A link from the person to the organisation node
@id56Lets other nodes in the graph reference them
description42Prose about the person, inside the markup
knowsAbout6The subjects they can speak to
honorificPrefix5Where the eight Dr. prefixes belonged
Neither url nor sameAs272Nothing an engine can disambiguate against
Properties present on the 518 Person nodes that carried a name, across the 1,079 home pages and 709 about pages read. Measured by Lantad on 6 October 2026. A node may carry several, so the rows do not sum.

The 162 names that appeared nowhere else on the page

Google's best practices ask that every author presented on the page is also in the markup. This corpus answers the reverse question, and it is the more interesting one. For each of the 518 named Person nodes, the delivered HTML was searched for the same name with every JSON-LD script element removed first. On 162 of them, across 57 hostnames, the name appeared nowhere in what was left. The structured data was the only place on that page where the person existed.

crayo.ai is the clearest case and it is not a defect. Its single JSON-LD block declares an Organization whose founder and employee arrays reference three identifiers, and the three Person nodes those resolve to carry, between them, name, url, jobTitle, description, image, knowsAbout, worksFor, memberOf and sameAs. Two of the three names appear nowhere in the 80,465 bytes of HTML the server delivered once that one block is removed: the document holds two occurrences of the first founder's name and both are inside it, and 73,916 bytes remain after it goes. A parser is told who runs the company in detail. Anything reading the delivered markup without that block finds one of the three. That is the inverse of the usual complaint about structured data, and on a small number of sites it is what is happening: the markup carries identity the rest of the page does not.

The commoner version is thinner. trueform.agency declares two Person nodes whose only property is a name, neither appearing in the delivered text; amplemarket.com declares an Organization whose founder array holds three Person nodes carrying nothing but a name; povio.com declares two people with jobTitle and sameAs on its home page and neither name appears in that page's delivered HTML outside the block. Measured against the readable text this scanner's extractor recovers rather than against the raw HTML, 228 of the 518 names were absent and 290 were present, a looser test that catches names sitting in an attribute or a script rather than in prose. The strict test is the 162.

Both numbers point at the same structural fact, and it is a prose parity problem wearing different clothes. This blog has measured pages where the schema said something the page never showed and pages where the answers in FAQ markup were not on the page. Those were defects, because the markup made a claim the page contradicted. A name is different: a person who exists only in the markup is not a contradiction, and an engine reading JSON-LD gets the information. The risk is narrower. If the only place a human appears is a block a renderer can drop, a template change can remove the entire human identity of a site without altering one visible word, and nothing in a browser will show it. That is an argument for checking what an AI crawler actually receives, which is what the GPTBot view tool renders for a single URL.

crayo.ai, 6 October 2026, LantadBot

  • GET / HTTP/1.1 200 OK, 80,465 bytes of HTML
  • script[type=application/ld+json] blocks 1
  • Organization founder 3 references by @id
  • Organization employee the same 3 references
  • Person nodes in the graph name, jobTitle, sameAs, knowsAbout, worksFor
  • search delivered HTML for the founder's name 2 matches
  • search again with the ld+json block removed 0 matches
  • bytes remaining after removal 73,916
One hostname's single JSON-LD block, and the result of searching the rest of the delivered HTML for the same names. Measured by Lantad on 6 October 2026. Identifiers are reproduced as delivered.

Education, government and healthcare named almost nobody

The corpus is stratified, which makes the sector comparison cheap to run and hard to dismiss. Education declared a person on 1 of 102 sites, the single case being unacademy.com. Government declared one on 1 of 86, nasa.gov. Healthcare managed 3 of 93: acibadem.com.tr, acponline.org and stjude.org. Against that, the SaaS stratum declared a person on 25 of 117 sites and the local media stratum on 8 of 30, the highest rate of the eighteen.

Part of that gap is not about people at all. Those three strata deploy less structured data of any kind: 30 of 102 education sites, 28 of 86 government sites and 44 of 93 healthcare sites carried a JSON-LD block that parsed, which is 29.4, 32.6 and 47.3 percent against nine of the eighteen strata at 60 percent or better and 98.1 percent in the Wix and Squarespace stratum. So the comparison was run again on the sites that had already declared an organisation, which removes the confound. Education: 1 of 25. Government: 1 of 22. Healthcare: 3 of 33. SaaS: 23 of 85. The gap narrows and it does not close.

What makes this worth a section rather than a row in a table is which sectors they are. A university's authority is its named academics. A hospital's is its named clinicians. A government department's is the named official responsible for a decision. Every one of those institutions publishes those names at length in prose, on staff directories and faculty pages this measurement never visited, and every one of them has a stronger claim to the experience and expertise half of what Google calls E-E-A-T than a three-person startup declaring its founders in JSON-LD. The startup is the one telling a machine about it.

The news stratum is the instructive exception, and only just. 8 of 60 news sites declared a person, with 404media.co and grist.org and spectator.co.uk among them, and news is the one sector where the byline is both an editorial convention and a documented structured data property. Even there, 52 of 60 did not. A related finding from this corpus is that 27 of 165 article pages carried a headline that was not the h1, which suggests article markup in this stratum is often a plugin's default rather than a considered description of the page. Sites that want to be named in an answer rather than merely read can start from the platform notes for ChatGPT and the equivalents for Perplexity and Google AI Overviews.

StratumSites readNamed a personOf those declaring an organisation
media-local3087 of 20, the highest rate in the corpus
saas1172523 of 85
news6087 of 47
finance9444 of 45
travel8043 of 27
healthcare9333 of 33
government8611 of 22, and it is nasa.gov
education10211 of 25, and it is unacademy.com
Sites declaring a Person node on the home page or the about page, by corpus stratum, with the same count restricted to the sites that already declared an Organization-family node. Measured by Lantad on 6 October 2026.

What this measurement does not show

The scan read two pages per site and nothing else, so every number above is a statement about a home page and one about page, never about a site. A university with a thousand faculty profiles carrying Person markup on each of them appears in this corpus as a site that named nobody, and that is the correct reading of what was measured rather than a flaw to be apologised for. The same limit applies to the 162 names that appeared only in markup: they appeared only in the markup of the page that was read.

No JavaScript was executed on either request. An earlier run over this corpus found that 27 of 404 home pages with no structured data in the HTML gained some once a browser ran the page, so a client-rendered Person node would be invisible here. The rate was 6.7 percent of the pages tested in that run, which is the right order of magnitude to hold against the 90 home page figure rather than a correction to apply to it.

Nothing here establishes that declaring Person markup changes whether an engine cites a site. This blog has published the measurement that cuts the other way: a citation audit in which structured data came third behind other factors, and the llms.txt study showing a file this product ships a tool for going almost entirely unread. Person markup is a signal that is cheap, documented, checkable and absent on about 90 percent of this corpus. That is the claim. Whether it moves a citation is a different measurement, and the honest position is that this scanner's own method notes do not score it today.

Finally, the about page was located by following the home page's own anchors, matching link text and paths against a fixed list of words such as about, team, leadership and company. 718 of 1,079 home pages offered such a link and 709 of those answered. The 361 that offered none are not sites without an about page; they are sites that do not link one from the home page in a form this method recognised, and a reader will notice that is itself a finding about internal links rather than about Person markup. Anyone who wants the same two checks on their own site can read what an AI crawler receives and look for a Person node in it.

  • Home page read as LantadBot 1,079 of 1,419 hostnames answered 200 with an HTML content type.
  • About page followed from the home page 718 home pages linked one, 709 answered 200 with HTML.
  • JSON-LD parsed with the scanner's own extractor 620 home pages carried at least one block that parsed.
  • JavaScript executed Neither request ran a browser, so a client-rendered Person node is not counted.
  • Interior pages, staff directories, author pages Two pages per site were read. A profile three clicks deep is outside the measurement.
  • Any effect on citations or rankings Nothing here measures whether the markup changes what an engine says.
What this run measured and what it did not, stated as the scan was configured. Measured by Lantad on 6 October 2026.

Written by

Lantad

Published .

Google published a question for site owners to ask themselves about their own pages, and it is not about markup: is it self-evident to your visitors who authored your content. The page it sits on was last updated on 5 October 2026, the day before this measurement ran, and it goes on to ask whether pages carry a byline where one might be expected and whether bylines lead to further information about the author. Those are questions about what a reader sees. This post measures the machine-readable half of the same question, because an answer engine deciding whether to repeat what your site says does not read your about page the way a person does. It reads what the page states in a form it can parse, and for a human being that form is one schema.org type.

Common questions

What is Person schema?

Person is the schema.org type for a human being, defined on schema.org as a person, alive, dead, undead or fictional. Published in JSON-LD on a page, it tells a machine that a named individual exists and, through properties such as jobTitle, worksFor, url and sameAs, what they do and which other records on the web describe the same individual. Lantad found it on 90 of the 1,079 corpus home pages read on 6 October 2026.

Does Google require Person schema?

No. Google's article structured data documentation, carrying Last updated 2026-09-08 UTC, states that there are no required properties and lists author as recommended. It does ask that all authors presented as authors on the page are also in the markup, that the Person type is used for people rather than the Organization or Thing type, and that author.name holds only the name, with a job title or an honorific prefix moved to jobTitle or honorificPrefix.

Is a name enough, or does a Person node need url or sameAs?

A name alone is a string that two people can share. Google's documentation says author.url is a link to a page that uniquely identifies the author and that Google can understand both sameAs and url when disambiguating authors. On the 518 named Person nodes Lantad measured on 6 October 2026, 193 carried url, 70 carried sameAs, 17 carried both and 272 carried neither.

Why did so few education and government sites declare a person?

The measurement does not say why, only that they did not on the two pages read. Education declared a person on 1 of 102 sites and government on 1 of 86 on 6 October 2026. Those strata also deploy less structured data overall, 30 of 102 and 28 of 86 carrying a JSON-LD block that parsed, but restricting the comparison to sites that already declared an organisation leaves the gap open at 1 of 25 and 1 of 22 against 23 of 85 in the SaaS stratum.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.