BlogFindings

Organization schema logo: 44 of 433 declared URLs returned no image, and 22 of those were 404

Lantad asked all 1,419 hostnames in this repository's committed corpus for robots.txt on 28 September 2026, read the home page of every host that allowed it with no JavaScript executed, then requested every logo URL the markup declared. 1,093 home pages answered HTTP 200 with HTML, 520 carried an Organization node, and 433 of those declared a logo. 389 of the 433 URLs returned image bytes: 22 answered 404, eight returned a web page, five never answered.

16 min read Lantad

The method was the same first two steps as the rest of this corpus work, plus a third. Every one of the 1,419 hostnames in this repository's two committed corpus seed files was asked for /robots.txt, the home page of every host that allowed this scanner at the root was requested once as LantadBot with redirects followed and no JavaScript executed, and then every logo URL those pages declared was requested in turn. 1,093 home pages answered HTTP 200 with an HTML content type, 635 carried at least one JSON-LD block that parsed, and 520 of those carried an Organization-family node, the structured data that names who is behind a page. 433 of the 520 declared a logo. The four requirements the declarations were then checked against are not this scanner's opinion: they are printed on Google's Organization structured data documentation, which carried Last updated 2026-09-08 UTC when it was read the same day.

In short

  • Lantad read the home page of 1,419 hostnames on 28 September 2026 and found 433 declaring an organization schema logo, of which 44 URLs returned no image at all: 22 answered 404, eight returned a web page, five never answered and one served a Windows icon file.
  • Google's Organization structured data documentation, carrying Last updated 2026-09-08 UTC, states that the logo image must be 112x112px at minimum, and 66 of the 261 raster logos Lantad fetched on 28 September 2026 were under that on at least one side, eight of them on both, among them cnn.com at 60 by 61 and cambridgeday.com at 16 by 16.
  • Google's documentation also requires the logo URL to be crawlable and indexable, and on 28 September 2026 eight of the 398 reachable logo URLs sat on a path their own robots.txt refuses Googlebot while seven were served with an X-Robots-Tag of noindex or none, 13 distinct sites between them.
  • 146 of the 433 declarations were ImageObject nodes carrying a width and a height, and on 16 of the 107 that could be checked against the bytes on 28 September 2026 the declared size was wrong, including atlassian.com declaring 512 by 512 for a file that is 96 by 96.
  • Nothing in this run measured whether any answer engine reads the logo property, so the 122 of 433 declarations that failed at least one published requirement are a defect in the markup rather than a measured loss of visibility.
StageSitesWhat happened
Hostnames asked1,419The committed corpus, an editorial frame rather than a random draw
Refused this crawler in robots.txt13Disallowed LantadBot at the site root, so no home page was requested
Did not answer 200 with HTML313194 answered 403, 92 never returned a status, 26 answered another code, and ramp.com answered 200 with markdown
Answered 200 with HTML1,093The readable set
Carried parseable JSON-LD635The denominator for anything about markup
Carried an Organization node520An identity a consumer can attach a logo to
Declared a logo on it43383.3 percent of the identity nodes
Logo URL returned image bytes38944 of the 433 returned something else or nothing
One GET of https://<host>/robots.txt and then https://<host>/ per hostname as LantadBot/1.0 (+https://lantad.co/bot), redirects followed, no JavaScript executed, then one GET of every declared logo URL. Measured by Lantad on 28 September 2026 across the 1,419 hostnames in this repository's two committed corpus seed files.

Did the declared URL return an image?

433 logo URLs were requested once each, as this scanner, with redirects followed and a twenty second timeout. 398 answered HTTP 200. Of the 35 that did not, 22 answered 404, five never returned a status at all, and eight returned a mixture of 400, 401, 403, 406 and 500. techcrunch.com declared a WordPress upload path ending resize=1200,1200 that answers 404 with a one byte body. blackrock.com declared an SVG under /blk-one-assets/ that answers 404 with 217 kilobytes of error page. Both were requested twice on the day and behaved identically.

The 398 that answered 200 are not 398 images. Eight of them returned HTML. nationalgeographic.com declares its logo as https://www.nationalgeographic.com/#logo, which is the home page with a fragment on the end, and the fetch duly returned 585 kilobytes of home page. altitudemarketing.com, fortnox.se, kempinski.com, mapfre.com, nav.no, projectant.io and sentry.io each returned a web page under the same property. One more, standardbank.co.za, returned a 410 kilobyte Windows icon file from its favicon directory, which is a real image and is not a format on Google's list.

That leaves 389 of the 433 returning image bytes, and 44 returning none. A 404 on a logo is a specific kind of rot: the markup was right when it was written and the file moved. Nothing on the page changes when that happens, no validator that reads only the document will see it, and the earlier run that checked whether the markup itself was well formed, which found 117 of 612 home pages carrying a schema defect, would pass every one of these 44 because the JSON parses and the type resolves. The same is true of the run that found 103 of 385 pages with JSON-LD naming no organization at all: a missing node and a broken pointer are different failures and only one of them is visible in the source. What a crawler receives, rather than what the page believes it published, is the thing this scanner's crawler view exists to show. The broader point is familiar from the images on these same pages, where 18,171 of 52,077 image elements carried an empty alt: the markup around an image is routinely written once and never checked again.

  • Image bytes returned 389 211 PNG, 128 SVG, 37 JPEG, 11 WebP, one AVIF and one GIF
  • HTTP 404 22 Including techcrunch.com, blackrock.com, metlife.com, redis.io and inquirer.com
  • Other non-200 13 Five never answered, and eight returned 400, 401, 403, 406 or 500
  • 200 with a web page 8 nationalgeographic.com declares the home page itself with a fragment on the end
  • 200 with an icon file 1 standardbank.co.za, 410KB of .ico, a format Google Images does not list
Outcome of one GET per declared logo URL, 433 requests, measured by Lantad on 28 September 2026 as LantadBot with redirects followed and a twenty second timeout.

Is the 112 pixel minimum a real constraint?

For raster files it is, and it bites more often than the 404s do. 261 of the 389 images were raster formats whose dimensions can be read out of the file header, and 66 of those 261 were under 112 pixels on at least one side. 58 of the 66 are the same shape: a horizontal wordmark that is wide enough and nowhere near tall enough, such as bankofamerica.com at 360 by 36, chubb.com at 550 by 92 and independent.co.uk at 504 by 60. Eight are small on both sides, and those are the ones worth naming, because a site does not usually choose them on purpose.

cnn.com declares https://media.cnn.com/api/v1/images/stellar/prod/cnnlogo.png?q=w_60,h_61, and the query string is the point: the site is asking its own image API for a 60 by 61 rendering, so the size is not an accident of an old file but a parameter somebody typed. cambridgeday.com declares a 16 by 16 PNG whose filename is cdfavicon.png. restofworld.org declares favicon-32x32-1.png. aig.com declares an 89 by 49 crop from an Adobe Experience Manager content fragment. atlassian.com and agentplace.io both declare 96 by 96, just under the line. The remaining 195 raster logos clear 112 on both sides, so this is a minority failure rather than a general one.

The 128 SVG files were not assessed against the pixel minimum, and that is a limit of this measurement rather than a pass. An SVG has no intrinsic pixel size in the sense the requirement means, and Google lists SVG among supported formats without saying how the 112 pixel floor applies to one. 87 of the 128 declare a viewBox or a width and height smaller than 112 on a side, but that number is not a failure count and this post does not use it as one. Reporting it as a failure would be the same error as scoring a design decision, which is what the AI visibility score avoids by keeping settings and measurements apart. The honest statement is narrower: 66 of 261 raster logos are under Google's published minimum, and the SVG half of the population is unassessed. That kind of split is also why the earlier count of 33 of 615 home pages declaring an aggregateRating separated what Google says it will show from what the markup merely contains, and why the run on 85 of 317 organizations naming a reference entry in sameAs counted destinations rather than judging them.

  • Google's published minimum 112px The floor on the Organization page
  • carpentercostin.com (100x70) 70px
  • trueform.agency (64x64) 64px
  • cnn.com (60x61) 60px Requested from its own image API as w_60,h_61
  • aig.com (89x49) 49px
  • restofworld.org (32x32) 32px Filename favicon-32x32-1.png
  • cambridgeday.com (16x16) 16px Filename cdfavicon.png
The eight raster logos under 112 pixels on both sides, shown by their shorter side in pixels, against Google's published minimum of 112. Measured by Lantad on 28 September 2026 by reading the dimensions out of the file header of each fetched image.

Crawlable and indexable: the two requirements a page cannot show you

Google's wording is that the image URL must be crawlable and indexable. Those are two separate controls in two separate places, and neither of them lives in the markup. Crawlable is decided by the robots.txt of whichever host serves the image, which is frequently not the host serving the page: 162 of the 398 reachable logos came from somewhere else, 24 of them from cdn.prod.website-files.com and the rest spread across Cloudinary, Contentful, Sanity, Ghost, WordPress.com and a long tail of one-offs. Indexable is decided by the X-Robots-Tag header on the image response, which Google's robots meta tag documentation defines for non-HTML files precisely because a PNG cannot carry a meta element.

Eight of the 398 logo URLs sit on a path their own robots.txt disallows to Googlebot: absa.co.za, atlassian.com, browserstack.com, eurostar.com, frontiersin.org, lemonde.fr, novavoice.app and wisesystems.com. Seven were served with an X-Robots-Tag naming noindex or none: asana.com, auth0.com, browserstack.com, frontiersin.org, hubspot.com, metoffice.gov.uk and pkobp.pl. browserstack.com and frontiersin.org appear on both lists, so 13 distinct sites publish a logo they have also told Google not to fetch or not to keep. browserstack.com is the clearest case, serving its logo from a WordPress host with the header noindex, nofollow, nosnippet, noarchive and disallowing the path as well.

The same robots.txt evaluation run against AI crawler tokens gives a larger number and a different meaning. 45 of the 398 logo URLs are disallowed to GPTBot and 42 to Google-Extended, and the list is dominated by news publishers: afr.com, aljazeera.com, corriere.it, faz.net, forbes.com, lemonde.fr, spiegel.de, theverge.com, usatoday.com and others. Those are deliberate policies rather than mistakes, and nothing in this run suggests otherwise. Checking a path against a file the way this comparison does is what the robots.txt tester exposes directly, and the token list it evaluates against is the same registry described on the crawler reference. The distinction between a rule aimed at an AI crawler and a rule that catches one on the way past is the recurring finding in this corpus, and it is the reason this scanner publishes its own bot's identity and behaviour rather than asking to be trusted.

Siterobots.txt allows GooglebotX-Robots-Tag on the image
browserstack.comNonoindex, nofollow, nosnippet, noarchive
frontiersin.orgNonone
absa.co.zaNoNot sent
atlassian.comNoNot sent
eurostar.comNoNot sent
lemonde.frNoNot sent
novavoice.appNoNot sent
wisesystems.comNoNot sent
asana.comYesnone
auth0.comYesnoindex
hubspot.comYesnone
metoffice.gov.ukYesnoindex
pkobp.plYesnoindex
The 13 sites whose declared logo URL fails one of Google's two access requirements, measured by Lantad on 28 September 2026 by evaluating the logo path against the serving host's robots.txt and reading the X-Robots-Tag on the image response.

What the markup claims about the image, and what the bytes say

227 of the 433 declarations were an ImageObject node, 203 were a bare string URL and three were an untyped object carrying a url property. Seven gave a value that is not an absolute URL: ally.com, eatfishwife.com and shein.com wrote protocol-relative addresses beginning with two slashes, while elderlawgroupwa.com, gosh.nhs.uk, groundcover.com and teamviewer.com wrote a path alone. All seven resolve against the page for a browser and all seven are outside the range schema.org gives the property, which is a reminder that a consumer is entitled to be stricter than a browser.

146 of the ImageObject nodes carried both a width and a height, which is the only place in this measurement where the markup makes a checkable claim about the file rather than merely pointing at it. 107 of those could be compared against a raster file that actually arrived. 91 matched the bytes exactly. 16 did not, and the direction of the error runs both ways: datadoghq.com declares 847 by 847 for a 2000 by 2000 file, thenationalnews.com declares 200 by 132 for one that is 2369 by 773, zendesk.com declares 512 by 512 for 1280 by 640, and wired.com declares 500 by 100 for 1201 by 631. Two of the 16 matter for the pixel minimum rather than for tidiness: atlassian.com and trueform.agency both declare 512 by 512 while serving 96 by 96 and 64 by 64 respectively, so markup that reads as compliant sits above a file that is not.

None of that is exotic. It is what happens when a template writes a number once and a build pipeline resizes the asset later, and it is the same class of drift as an identifier that no longer points anywhere, which is why the run on 313 of 615 home pages giving no schema node an @id treated graph wiring as its own failure mode. The fix is boring and it is the same on every stack: generate the property from the asset rather than typing it beside the asset, which for a framework build is the pattern set out in the Next.js guide.

SiteDeclared in markupRead from the file
thenationalnews.com200 x 1322369 x 773
metalbear.com512 x 5121977 x 1007
wired.com500 x 1001201 x 631
zendesk.com512 x 5121280 x 640
datadoghq.com847 x 8472000 x 2000
atlassian.com512 x 51296 x 96
trueform.agency512 x 51264 x 64
fidelity.com600 x 60164 x 36
The eight largest disagreements between the width and height an ImageObject declared and the dimensions read from the file it pointed at, of 16 mismatches in 107 comparable pairs. Measured by Lantad on 28 September 2026.

What this measurement does not show

311 of the 433 declarations passed every check applied here and 122 failed at least one. That second number is a count of published requirements broken, and it is not a count of anything lost. No citation, ranking, knowledge panel or answer engine output was measured in this run, and nothing here supports a claim that fixing a logo URL gains a site anything. Google's own wording is that the property can help it understand which logo to show. That is a statement about a search feature, not a promise, and this post carries it no further than the source does.

The gap between the two is where the honest limit sits. This scanner weights logo at 15 points inside its entity signal set, and that weight is a decision taken from a frequency count across a small research sample, not a measurement of what any engine rewards. Publishing the number beside the finding is the point: a reader can see that the scoring is a judgement and the 44 broken URLs are not. The same separation runs through the write-up of what a generative engine optimization programme can and cannot claim, and through the survey work collected on the research page.

Five other limits are worth stating plainly. Only JSON-LD was read, so a logo expressed in microdata or RDFa is invisible here. No JavaScript was executed, so markup injected in the browser counts as absent, which is a statement about what a crawler is served rather than about what a person sees. One page was read per hostname, so every figure describes home pages rather than sites. Every request was one attempt from one network location, so a briefly throttled host is undercounted rather than counted. And the corpus is an editorial sampling frame assembled for platform and industry coverage, not a random draw of the web, so each rate supports a statement about these 1,419 hostnames and nothing wider. Readers arriving from the question of how content gets cited in Google AI Overviews should take the same care: a logo is an identity signal with one documented consumer, and a documented consumer is not evidence of an effect.

The four gates every declared logo was put through, in order, with the count each one caught. Measured by Lantad on 28 September 2026 across 433 declarations on 1,093 readable corpus home pages.

Written by

Lantad

Published .

An organization schema logo is one of the few pieces of structured data on a home page that points at something outside the document. Most of what a page declares about itself is text sitting in the same bytes a crawler already has: a name, a description, a language tag, a set of profile addresses. A logo is an address, and an address is a promise that something is there. This run went and checked, because a declaration and a resource are different things and only one of them survives a redesign.

Common questions

Is the logo property required in Organization structured data?

No. Google's Organization structured data documentation, carrying Last updated 2026-09-08 UTC when it was read on 28 September 2026, states that there are no required properties and that you should add the properties that apply to your organization. logo appears among the recommended ones, with the note that it can help Google understand which logo to show, for example in Search results and knowledge panels. schema.org defines it simply as an associated logo, with an expected type of ImageObject or URL at release V30.1 dated 2026-09-16.

What size does Google require an Organization logo to be?

At least 112 by 112 pixels, which is the minimum Google's Organization page states. In this corpus 66 of the 261 raster logos measured on 28 September 2026 were under that on at least one side, and 58 of the 66 were wide horizontal wordmarks that fail only on height. Google does not say how the minimum applies to an SVG, so the 128 SVG logos in this run were not assessed against it and are not counted in that 66.

Can a logo URL be blocked by robots.txt?

Yes, and eight of the 398 reachable logo URLs in this corpus were, measured on 28 September 2026. The robots.txt that decides it belongs to whichever host serves the image, which was a different host from the page on 162 of the 398. A separate seven sites sent an X-Robots-Tag of noindex or none on the image response, which is the indexing half of the same requirement. Neither control is visible in the page markup, so neither shows up when a developer checks the logo by loading it in a browser.

Does fixing the logo property improve AI visibility?

This measurement does not show that, and no figure in this post supports it. The run counted declarations, fetched the files and checked them against published requirements. It measured no citation, no ranking and no answer engine output. Google documents a consumer for the property in Search results and knowledge panels, which is more than most schema properties have, but a documented consumer is evidence that something reads the field rather than evidence of an outcome from filling it in.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.