BlogFindings
ImageObject schema: 253 home pages declared 943 nodes, and 685 carried nothing but a URL and a size
Lantad asked all 1,419 hostnames in this repository's committed corpus for robots.txt on 8 October 2026 and read each home page it was allowed to read, with no JavaScript executed. 1,084 answered HTTP 200 with HTML, 618 of those carried parseable JSON-LD, and 253 declared ImageObject schema across 943 nodes. 685 of the 943 carried only an image URL and pixel dimensions, 128 carried a caption, and 33 nodes on 2 sites carried the pair of properties Google's image metadata documentation requires.
This run read one page per hostname across the whole committed corpus and counted what those nodes actually contain. The short answer is that ImageObject is common and nearly empty: 253 sites declared it, 943 nodes between them, and 685 of those nodes carried an address and a pixel size and no other information of any kind. The longer answer is more interesting than a scolding, because 423 of the 943 were describing a company logo, and a logo genuinely does not need a caption or a licence. The gap narrows when you ask what each node was for. It does not close.
In short
- ImageObject schema was declared by 253 of the 1,084 home pages that served Lantad an HTML page on 8 October 2026, carrying 943 nodes between them, which ranks it third of the 147 schema.org types found on those pages.
- 685 of those 943 nodes carried nothing beyond an image URL and pixel dimensions, with no caption, no credit line, no copyright notice and no licence.
- Google's image metadata documentation, carrying Last updated 2025-12-10 UTC, requires contentUrl plus one of creator, creditText, copyrightNotice or license; 33 nodes on 2 of the 253 sites carried that pair.
- None of the 943 nodes named a creator and none carried acquireLicensePage, the two properties that describe who made the image and where it can be licensed.
- 423 of the 943 nodes were the value of a logo property rather than a description of page content, and of the 520 that were not logos, 33 carried any of the four properties Google asks for.
What does ImageObject schema do for an AI crawler?
It supplies the words a photograph does not contain. An engine assembling an answer has to decide what a picture shows, whether it illustrates the claim it sits beside, and whether it may be reproduced. None of that is in the image file. It is in the alt attribute, the surrounding prose, and the ImageObject node, and of those three the node is the only one written as machine readable fields with defined meanings.
The measurement here follows the same route every corpus run on this site takes, and it is written down in full on the methodology page. Each of the 1,419 hostnames in the two committed seed files was asked for its robots.txt as LantadBot, the identified crawler described at our bot page. 1,110 returned HTTP 200 and 1,077 of those parsed as plain text. 12 sites closed the site root to LantadBot in their own file, so no page was requested from them. Of the 1,407 home pages that were requested, 216 answered HTTP 403, 53 answered HTTP 503, 32 never returned a status at all, 22 answered some other status, and 1,084 answered HTTP 200 with an HTML content type. Those 1,084 are the denominator for everything below.
618 of the 1,084 carried at least one JSON-LD block that parsed as JSON. 14 pages served a block that did not parse, and whatever was inside those is missing from every figure here. Across the 618 pages the parser found 147 distinct schema.org types. ImageObject came third, on 253 sites, behind Organization on 412 and WebSite on 410 and ahead of SearchAction on 224 and PostalAddress on 198. That ordering is worth holding onto: ImageObject is not an exotic type that a handful of publishers bother with, and it is far more widely declared than the types this blog has found nearly absent, such as Service, on 20 of 616 pages. It sits in the small group of types that ordinary sites really ship, which is also the group schema.org's own domain counts put at the top of the vocabulary.
Flow: 1,419 hostnames to robots.txt asked; robots.txt asked (disallowed) to 12 closed the root; robots.txt asked (allowed) to 1,407 pages requested; 1,407 pages requested (403, 503, no status) to 323 served no page; 1,407 pages requested to 1,084 answered 200 HTML; 1,084 answered 200 HTML to 618 carried JSON-LD; 618 carried JSON-LD to 253 declared ImageObject.
685 of 943 nodes carried nothing but a URL and a size
Counting properties across all 943 nodes gives a clear shape. 875 carried url and 176 carried contentUrl, with 126 carrying both, so 925 nodes named an image file somewhere and 18 named none at all. 694 carried both width and height. After that the counts fall away: 128 carried a caption, 39 a name, 26 a description, 63 a credit line, 33 a copyright notice and 24 a licence.
685 of the 943, which is 72.6 percent, carried no property at all beyond url, contentUrl, width and height. Read as a sentence, such a node says there is an image at this address and it is this many pixels across. An engine reading it learns the file exists and nothing about what it shows. The pixel dimensions are the one field almost always present and the one field a reader of an answer never needs.
The 18 nodes that named no image are the sharpest version of the same problem. All 18 carried only @context and @type, which is to say they declare that an image object exists and then stop. They sit on three sites, rijksoverheid.nl, mskcc.org and nationalgeographic.com. A further 19 nodes gave a relative path such as /static/images/logo.svg rather than an absolute URL, and one gave an array where a string was expected. None of those are fatal to a parser that resolves relative URLs against the page, but they are the same class of defect this blog found in 117 of 612 pages carrying a schema markup defect.
One oddity is worth recording because it says something about how these nodes are generated. 139 of the 943 carried inLanguage, a property that describes the language of a work. An image has no language. The same property was absent from 429 of the 615 pages carrying parseable JSON-LD in a run on 26 September 2026, which suggests a generator filling fields it can compute rather than fields that carry information. The caption, the one field that would tell an engine what the picture shows, appeared on 128 nodes across 93 sites, and the sampled values read like alt text written for a logo: company names, and occasionally something useful such as a dashboard screenshot described in a clause. On the same corpus, 18,171 of 52,077 image elements carried an empty alt attribute, so the absence of a written description is consistent across both places a page could put one.
| Property | Nodes | Share | What it tells an engine |
|---|---|---|---|
| url | 875 | 92.8% | Where the file is, loosely |
| width and height | 694 | 73.6% | Pixel size, which no answer needs |
| contentUrl | 176 | 18.7% | Where the actual bytes are |
| inLanguage | 139 | 14.7% | The language of an image, which has none |
| caption | 128 | 13.6% | What the picture shows |
| creditText | 63 | 6.7% | Who is credited |
| name | 39 | 4.1% | A short label |
| copyrightNotice | 33 | 3.5% | Who owns it |
| description | 26 | 2.8% | What the picture shows, at length |
| license | 24 | 2.5% | The terms of use |
| creator | 0 | 0% | Who made it |
| acquireLicensePage | 0 | 0% | Where to license it |
What Google requires in ImageObject schema, and how many met it
There is a published bar to measure against, so this does not have to be a matter of opinion. Google's image metadata documentation, carrying Last updated 2025-12-10 UTC when it was read on 8 October 2026, sets out the structured data for licensable images. Under Required properties it names contentUrl, described as a URL to the actual image content, and adds that Google also supports the url property if contentUrl is absent while recommending contentUrl instead because url is not as precise. Then it requires one of four more: creator, creditText, copyrightNotice or license. The page says that once one of those four is present the other three become recommended in the Rich Results Test. acquireLicensePage, which names the page where an image can be licensed, is listed under Recommended properties rather than required.
Against that rule the corpus reads as follows. 63 of the 943 nodes carried at least one of the four, and those 63 sit on 3 of the 253 sites: abc.net.au with 30, kempinski.com with 24 and rijksoverheid.nl with 9. Narrowing to the required pair, 33 nodes on 2 sites carried contentUrl together with one of the four, those sites being kempinski.com and rijksoverheid.nl. Allowing the url fallback Google says it still supports lifts the figure only back to 63, because every node carrying one of the four already named a file one way or the other.
Two of the four never appeared at all. Not one of the 943 nodes carried creator, so no page in this corpus named the photographer, illustrator or agency behind any image it marked up. Not one carried acquireLicensePage. The 24 licence values that did exist were all absolute URLs, which is the shape the documentation asks for, and all 24 belonged to a single site pointing at one legal page.
It is worth being exact about what this does and does not imply. The documentation is Google's, and the feature it describes is the licence information shown in Google Images, not a statement about what any answer engine does. Nothing measured here shows that an engine read these properties, ignored them, or would have cited a page differently had they been present. What the figures support is narrower and still useful: a site that wanted the licensable image treatment would not get it, because the markup does not meet the published requirement. For the separate question of what gets quoted in an AI answer, the guidance this site keeps is on getting cited in Google AI Overviews, and image licensing is not part of it.
- contentUrl, required 176 of 943 nodes. Google supports url as a less precise fallback, present on 875.
- One of creator, creditText, copyrightNotice, license, required 63 of 943 nodes, on 3 of the 253 sites.
- contentUrl plus one of the four, the required pair 33 nodes on 2 sites: kempinski.com and rijksoverheid.nl.
- creator, recommended once one of the four is present 0 of 943. No page named who made the image.
- acquireLicensePage, recommended 0 of 943. No page said where an image could be licensed.
Why contentUrl appeared on 176 of the 943 nodes
The two properties are not synonyms and the difference is in the vocabulary itself. Schema.org's definition of ImageObject describes contentUrl as the actual bytes of the media object, for example the image file or video file. url, inherited from Thing, is the address of the item being described, which for an image object may reasonably be the page the image appears on rather than the file. A node carrying only url is therefore ambiguous by construction: a parser cannot tell from the markup alone whether it has been handed a picture or a web page that contains one.
That ambiguity is why Google's documentation says it uses contentUrl to determine which image the photo metadata applies to, and recommends it over url while continuing to accept url because existing markup uses it. In this corpus existing markup overwhelmingly uses it. 875 nodes carried url against 176 carrying contentUrl, and 126 carried both, which leaves 50 nodes relying on contentUrl alone and 749 relying on url alone.
The pattern holds for the other media type this blog has measured. Of the 43 VideoObject nodes found on this corpus in September, all carried the three properties Google's video documentation requires, which is a better result than anything here and probably reflects that video markup is written deliberately by people chasing a video result, while image markup arrives from a template. The same template effect explains the dimensions: width and height on 694 nodes is the signature of a content management system that knows the file it stored and nothing about the picture inside it.
Format matters less than it used to. Everything above was parsed from JSON-LD, which is where this corpus keeps its structured data: an earlier run over the platform seed file found microdata on 22 of the 385 pages that answered, so a node expressed only in attributes is a small enough population to ignore without distorting the counts. It is still a gap, and an ImageObject written as microdata is invisible to every figure in this post.
contentUrl, on 176 of 943
- Defined as the actual bytes of the media object
- Unambiguous: it names the file
- The property Google uses to decide which image metadata applies to
- 50 nodes carried it without url
url, on 875 of 943
- Inherited from Thing: the address of the item described
- May be the file or the page holding it
- Accepted by Google as a less precise fallback
- 749 nodes carried it without contentUrl
423 of the nodes were the logo, not the picture
A count of empty nodes is only damning if the nodes were supposed to describe photographs. So this run also recorded where each node sat in the document, by the property whose value it was. The answer reframes the finding. 423 of the 943 were the value of a logo property, 343 were the value of an image property, 146 stood at the top level of their own block, 18 were a primaryImageOfPage, 10 an associatedMedia and 3 a screenshot. By the type of the node holding them, 216 hung off an Organization and 177 off a NewsMediaOrganization.
A logo does not need a caption, and a company does not need to tell anyone the licence terms of its own wordmark. On the 423 logo nodes, 84 carried a caption and 30 carried one of Google's four properties, and the remaining sparseness is defensible. This is the same population as an earlier run that resolved declared logo URLs and found 44 of 433 returning no image at all, which is a real defect in a way that a missing licence on a logo is not.
Take the logos out and the finding survives. 520 nodes were not logos, spread across 116 sites, and of those 520 just 33 carried any of the four properties and 44 carried a caption. The 343 nodes sitting under an image property are the starkest group: they are on only 43 sites, not one of them carried any of the four, and 17 carried a caption. These are the nodes most likely to be describing actual page content, and they are the emptiest. 386 of the 520 non-logo nodes carried nothing but a URL and a size.
That matters more for the entity question than for the licensing one. A logo that is marked up, resolvable and captioned is a signal about who publishes a page, which is the input behind the confidence an engine has in a brand, and this blog has measured the same signal going missing when a wordmark is drawn as unlabelled vector paths with no text. Named people behave the same way: 111 of 1,079 sites named a human in their markup, and 155 of the ImageObject nodes here hung off a Person. The pattern across all of it is that sites declare the existence of a thing and skip the description.
What this run measured, and what it did not
One home page carried 449 of the 943 nodes. eltiempo.com, which sits in the corpus's news stratum, marks up every article image on its front page, and its 449 nodes are 47.6 percent of the total on their own. The next largest are abc.net.au with 31, scmp.com with 30, kempinski.com with 25 and rijksoverheid.nl with 23. 150 of the 253 sites declared exactly one node and the median site declared one. Any node count in this post is therefore a count of one news home page plus a long tail, and it should be read that way.
The reassuring part is that removing the outlier changes no conclusion. All 449 of the eltiempo.com nodes are in the bare group, carrying a URL and a size and nothing else. Excluding the site leaves 494 nodes, of which 236 are bare, while the caption count stays at 128 and the count carrying one of Google's four stays at 63, because eltiempo.com contributed none of either. The proportion of bare nodes falls from 72.6 percent to 47.8 percent and every metadata finding above holds unchanged.
The rest of the limits are the ones every run on this corpus carries, and they are set out alongside the crawlability study. One page was read per hostname, the home page, so a site that captions images properly on its article pages and not on its front page counts as sparse here. No JavaScript was executed, so a node injected by a script was not seen, and an earlier run measured that gap at 27 of 404 pages gaining structured data when a browser ran the page. The 335 hostnames that served no page are absent from every figure, and the 216 that answered HTTP 403 are the sites most likely to hold firm views about crawler access, so these proportions describe sites that admit an identified crawler. The corpus is an editorial sampling frame built for platform and industry coverage, not a random draw, so every number describes these 1,419 hostnames on one date and nothing wider.
Two gaps are our own rather than the corpus's. Lantad does not validate ImageObject: the requirements table in core/src/schema.ts names Organization, Product, Article, FAQPage and BreadcrumbList, so an ImageObject carrying nothing but a URL raises nothing in a Lantad report, and a reader checking their own AI visibility would not learn it from us today. And no engine was observed in this run at all. Every request came from this scanner, from one network location, so nothing here is evidence that any crawler read, used or ignored an ImageObject node. If you want to see what one crawler receives from your own pages rather than what this corpus declared, what GPTBot sees fetches it live. What a markup audit cannot tell you is whether the image itself arrives, which is a separate measurement: 3,155 of 51,990 images on this corpus carried no URL in the HTML at all.
-
Declaration countsMeasured 253 sites, 943 nodes, property presence read from the delivered bytes of one page each. -
Against Google's ruleMeasured The required pair was read from Google's own documentation on the same day and applied to every node. -
Engine behaviourNot measured No crawler was observed. Nothing here shows an engine read or ignored these properties. -
Client rendered nodesNot measured No JavaScript was executed, so a node added by a script is missing from every count. -
Pages other than the home pageNot measured One page per hostname, so article level captioning is invisible here. -
Lantad's own coverageGap ImageObject is not in the requirements table in core/src/schema.ts, so an empty node raises nothing in a report.
Lantad
Published .
ImageObject schema is the part of structured data that describes a picture rather than a page: which file it is, how large it is, what it shows, who made it and on what terms it can be used. It matters to an AI crawler for a specific reason. The crawler downloads bytes, and the bytes of a photograph tell it almost nothing. The surrounding markup is the only place the page can say in words what the picture is.
Common questions
What is ImageObject schema?
ImageObject is the schema.org type that describes an image as a thing in its own right rather than as an attribute of a page. It carries contentUrl for the file itself, width and height for its dimensions, caption and description for what it shows, and creator, creditText, copyrightNotice and license for who made it and on what terms it can be used. It was the third most declared type on the 618 pages carrying JSON-LD in this corpus on 8 October 2026, on 253 sites.
Which ImageObject properties does Google require?
contentUrl, plus one of creator, creditText, copyrightNotice or license. Google's image metadata documentation, carrying Last updated 2025-12-10 UTC, states that it uses contentUrl to determine which image the metadata applies to and that it also supports url as a less precise fallback. acquireLicensePage is recommended rather than required. Of the 943 nodes measured on 8 October 2026, 33 on 2 sites carried the required pair.
Does ImageObject schema help a site get cited by an AI answer engine?
This run does not answer that, and no figure in it should be read as though it did. Every request came from Lantad's own crawler, so nothing was observed about engine behaviour. What can be said is that an image file carries no words, so the caption, the description and the surrounding prose are the only places a page states what a picture shows, and 685 of the 943 nodes measured here supplied none of them.
Is an ImageObject node with only a URL and dimensions a defect?
Not necessarily. 423 of the 943 nodes were the value of a logo property, and a company logo needs no caption or licence terms, so sparseness there is reasonable. The finding is firmer for the 520 nodes that were not logos, where 33 carried any of the four properties Google asks for and 386 carried nothing but a URL and a size, and firmest for the 18 nodes that carried only @context and @type and named no image at all.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.