BlogFindings
Open Graph tags: 370 of 948 home pages left out one of the four properties the specification requires
Lantad requested the home page of all 1,419 hostnames in this repository's committed corpus on 29 September 2026 and read the head of every response with no JavaScript executed. 1,135 answered HTTP 200 with an HTML content type, and 948 of those carried at least one og: property. Of those 948, 578 carried all four properties the Open Graph protocol calls required and 370 did not. On 14 of the 767 pages declaring an og:url, the URL the page gave as its own permanent identifier did not point at the page that served it.
This run checked what is actually in that block. Lantad requested the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026, sent as its own declared crawler user agent with redirects followed and a twenty second timeout, and parsed the head of the bytes that came back with no JavaScript executed. 1,135 answered HTTP 200 with an HTML content type. The naming half of this block has been measured here before and is not the subject: a run four days earlier found that 38 of 1,083 home pages carried neither a usable title element nor an og:title, and that where both existed they disagreed often. This run is about everything under the title. What type does the page claim to be, what address does it give as its own, and is any of it written where a parser will find it.
In short
- Lantad read the head of 1,135 home pages on 29 September 2026 and found Open Graph tags on 948 of them, but only 578 of those 948 carried all four of og:title, og:type, og:image and og:url, which the Open Graph protocol at ogp.me names as the four required properties for every page.
- On 14 of the 767 pages that declared an og:url, that URL did not resolve to the page that served it: six pointed at a different registrable host, among them an RFC 1918 private address at quebec.ca, a vendor-internal hostname at mountsinai.org, a hosting provider's staging hostname at browserstack.com and the literal template placeholder https://www.your-domain.com/your-page.html at hospitalitaliano.org.ar, and eight were not absolute URLs at all.
- 27 of the 1,135 pages wrote Open Graph names into a name attribute rather than the property attribute the specification uses, and on 12 of those there was no property-based og:title anywhere in the head, so dropbox.com and cedars-sinai.org shipped a complete Open Graph block that a parser following the specification does not see.
- Google's title link documentation, carrying Last updated 2025-12-10 UTC, names og:title among the nine sources it uses to choose a title link, and it is the only published vendor documentation this run could find that commits to reading any part of the Open Graph block.
- 25 of the 776 og:type values were not one of the thirteen global types the specification defines, including the strings illinois-home-page, BECU:home, webiste, hospi and object.
| Stage | Count | What happened |
|---|---|---|
| Hostnames requested | 1,419 | The committed corpus, an editorial frame rather than a random draw |
| Answered 200 with an HTML content type | 1,135 | The denominator for the first figure only |
| Carried at least one og: property | 948 | 83.5 percent of the pages read |
| Carried all four required properties | 578 | 61.0 percent of the 948 |
| Missing at least one required property | 370 | 39.0 percent of the 948 |
| Declared an og:url | 767 | Of which 14 did not point at the page that served them |
| Declared an og:type | 776 | Of which 25 were not a type the specification defines |
| Wrote og: names into a name attribute | 27 | 12 of them had no property-based og:title at all |
Do AI crawlers read Open Graph tags?
One publisher documents an answer in enough detail to quote, and it is Google. Google's documentation on influencing your title links, carrying Last updated 2025-12-10 UTC, lists nine sources it uses to choose the title link for a result, and the fourth of them is content in og:title meta tags. That is a commitment about one property, on one surface, from one vendor. It is also the whole of the published evidence. No crawler operator behind the tokens this scanner evaluates documents what it does with og:type, og:url or og:site_name, and none of them documents reading the Open Graph block at all.
So the honest framing of this measurement is narrow and it is worth stating before any figure. Nothing here shows that an AI crawler read these tags, because no crawler was observed reading anything: this run fetched pages, it did not inspect anybody's retrieval pipeline. What it shows is what a parser following the Open Graph protocol would find if it looked, and how often what it would find is wrong, absent or unreachable. Whether a given engine looks is undocumented, and undocumented is the finding rather than a gap in the method.
What can be said without any vendor's help is that this block is frequently the only machine-readable description a page carries. 519 of the 1,135 pages read here carried no parseable JSON-LD script block at all, and 339 of those 519 did carry an og:title. For that group the Open Graph block is not a second opinion sitting next to structured data. It is the only structured statement the page makes about itself. That is consistent with the earlier finding that 141 of 382 home pages carried no structured data at all in the raw HTML, measured on a smaller slice of the same frame. The remaining 149 of the 519 carried neither a parseable JSON-LD block nor a single og: property, which is the floor case: a page whose only self-description is its title element and whatever prose survives.
The specification itself is short and unambiguous, and it is worth reading at the source rather than from a summary. It lives at ogp.me, which is not a host this site links to, so the address is given here as plain text. It defines four required properties for every page, og:title, og:type, og:image and og:url, and it defines thirteen global types. The three sections below count each of those against what the corpus actually served.
- Google names og:title as a title link source One of nine sources listed in Google's title link documentation, Last updated 2025-12-10 UTC. The only vendor commitment this run found about any og: property.
- og:title, og:type, og:image, og:url are required Stated by the Open Graph protocol specification at ogp.me. 578 of the 948 pages carrying any og: property carried all four.
- The specification uses the property attribute Every example in the specification is written as property="og:title". 27 pages in this corpus used a name attribute instead.
- An AI crawler reads og:type or og:url No operator documentation behind the crawler tokens this scanner evaluates names either property. Undocumented, not disproved.
- An AI answer engine prefers og:title over the title element Google publishes the list of candidates and no weighting. Nothing observable from outside an engine settles it.
- This run observed a crawler reading these tags It did not. One GET per host from one network location, no server logs held, no retrieval pipeline inspected.
370 of 948 pages left out a property the specification requires
948 of the 1,135 pages carried at least one og: property. Counting the four the specification calls required, 578 of those 948 carried all four and 370 carried three or fewer. The absences are not evenly spread across the four. og:url was missing on 181 of the 948, og:image on 156, og:type on 172 and og:title on 51. So the property most often left out is the one that identifies the page, and the property most often present is the one that names it, which is the order you would expect from a block whose surviving purpose is drawing a card: a card needs a headline and a picture, and it does not visibly break when the identifier is absent.
That asymmetry is the practical point of this section. A missing og:image degrades a preview and somebody notices, because the preview is the thing people look at. A missing og:url degrades nothing anybody sees, so it survives indefinitely. The same logic explains why 365 of the 948 carried no og:site_name: the site's name is the part a card renders from the domain anyway, so leaving it out costs nothing visible while removing the one property in the block that states, in a field designed for it, what the publisher is called.
Presence also varies sharply by the kind of site, and the split follows the content management system rather than the sector. 53 of the 54 Wix and Squarespace pages that answered carried og:title, as did 111 of 118 SaaS pages, 55 of 60 news pages and 30 of the 31 Framer pages. Against that, 27 of the 41 WordPress small business pages carried any og: property at all, 26 of 31 static documentation sites did, and 50 of 71 ecommerce pages did. The platforms that generate the head for you generate this block completely; the stacks where somebody hand-assembled the head are where properties go missing. That is the same pattern the logo declared in Organization markup followed, and the same one behind the early five page finding that 85 meta elements in five page heads included two that addressed a crawler.
None of that makes an incomplete block a scored defect here. Lantad does not weight og:type in a grade and this post is not an argument that it should; what the scanner does and does not count is set out in the scoring methodology. The claim is narrower and it is about drift: a block that four out of ten sites cannot keep complete is a block whose contents nobody is checking, and the next three sections are what turns up when somebody does.
14 og:url values did not point at the page that served them
The specification defines og:url as the canonical URL of the object, used as its permanent identifier. It is the closest thing the block has to a primary key: the string by which anything consuming the page is invited to recognise it later. 767 of the 1,135 pages declared one. On 14 of those 767 the declared URL did not resolve to the page that served it, and the fourteen split into two kinds that fail for different reasons.
Six declared an absolute URL on a different registrable host from the one that answered the request. Four of the six are hostnames that were never meant to leave a private network or a build pipeline. quebec.ca, the government of Quebec, declares an og:url of https://10.3.3.229/, an address in the RFC 1918 private range that resolves to nothing on the public internet. mountsinai.org declares https://lit-mshsclh-p02.mshs.cloud.opentext.com/, which is a content management vendor's internal instance hostname. browserstack.com declares https://bstackprod.kinsta.cloud/, its hosting provider's origin hostname rather than its own domain. blogbowl.io declares https://ready-knowledge-301044.framer.app/, the preview URL a Framer site carries before a custom domain is attached. A fifth, alsea.com.mx, declares https://www.alsea.net/, a real and related corporate domain that is nonetheless not the host that served the bytes. The sixth is the one that needs no interpretation: hospitalitaliano.org.ar, a large Buenos Aires hospital, declares an og:url of https://www.your-domain.com/your-page.html, the literal placeholder from a template that nobody filled in. The same page declares an og:type of hospi.
The other eight are not absolute URLs at all, which the specification's definition of a permanent identifier does not allow. Five are effectively empty: washington.edu, rte.ie and 1password.com each declare an og:url of a single forward slash, southafrica.net declares /us/en/, and med.or.jp declares a protocol relative //www.med.or.jp/. Two declare a bare hostname with no scheme, uba.ar as uba.ar and trip.com as www.trip.com. The last, acc.org, declares http%3a%2f%2fwww.acc.org%2f, which is its own URL percent encoded once too often by whatever wrote the tag.
Ten of the fourteen sit in the government, education and healthcare strata. That is the same constituency that runs the largest and longest lived content management systems, and it is worth being precise about what the failure costs, because it is easy to overstate. A wrong og:url does not stop a page being fetched, indexed or quoted. What it does is offer a machine a false identifier for the thing it just read, next to a rel=canonical that almost always says something different: an earlier run found that canonical tags on five of five pages pointed at themselves. A page that answers the question "what are you" twice and differently is the same problem measured elsewhere as 412 of 1,080 home pages naming their site in neither source Google reads first, and it is the mechanism behind weak entity confidence: not an absence of signal, but two signals that cannot both be true.
| Hostname | Stratum | Defect | og:url as returned |
|---|---|---|---|
| quebec.ca | government | Different host | https://10.3.3.229/ |
| hospitalitaliano.org.ar | healthcare | Different host | https://www.your-domain.com/your-page.html |
| mountsinai.org | healthcare | Different host | https://lit-mshsclh-p02.mshs.cloud.opentext.com/ |
| browserstack.com | saas | Different host | https://bstackprod.kinsta.cloud/ |
| alsea.com.mx | travel | Different host | https://www.alsea.net/ |
| blogbowl.io | spa-startups | Different host | https://ready-knowledge-301044.framer.app/ |
| acc.org | healthcare | Not absolute | http%3a%2f%2fwww.acc.org%2f |
| med.or.jp | healthcare | Not absolute | //www.med.or.jp/ |
| uba.ar | education | Not absolute | uba.ar |
| washington.edu | education | Not absolute | / |
| rte.ie | news | Not absolute | / |
| 1password.com | saas | Not absolute | / |
| trip.com | travel | Not absolute | www.trip.com |
| southafrica.net | travel | Not absolute | /us/en/ |
25 og:type values were not a type the specification defines
776 pages declared an og:type. The specification defines thirteen global types: website, article, book, profile, payment.link, four music types and four video types. 751 of the 776 declared one of those thirteen, and the distribution is what a corpus of home pages should produce, 693 website and 58 article. The other 25 declared something that is not a type at all.
The invalid values divide into two groups, and the first is an honest guess. Three pages declared product and one declared product.group, which are real types in Facebook's own extended vocabulary but not in the global set the specification lists, so they are reasonable strings written by somebody who knew the block existed. The second group is a template leaking. Twelve pages declared a word describing the page rather than a type of object: Homepage twice, homepage once, Page twice, page twice, Home Page variants, Page d'accueil twice, site once and web once. Six more declared a fragment of something internal. illinois-home-page is a page identifier. BECU:home is a namespaced content key from a credit union's content management system. Web Area Homepage is a section label. Education and Articles and Company are categories. object and Object are the word from the specification's own prose, copied into the field it describes. One, webiste, is website misspelled, and one, hospi, is the same truncated string that hospitalitaliano.org.ar put in its og:url placeholder.
What these have in common is that every one of them is a field bound to the wrong source. Somebody wired og:type to a content type, a section name or a page title in a template, and because nothing renders it and no build step checks it, the wrong binding shipped and stayed. That is the identical shape as 117 of 612 home pages carrying a JSON-LD defect and as 27 of 165 article pages declaring a headline the page's own h1 does not contain. Markup that no human reads is markup that drifts from whatever generates it, and the drift is invisible until something parses it.
Whether an invalid og:type costs anything is a separate question and the honest answer is that nobody publishes one. A consumer that does not recognise the value will most likely fall back to treating the page as a website, which for 24 of these 25 home pages is the right answer anyway. The finding is not that these pages are penalised. It is that a quarter of the pages in this corpus declaring a type could not be checked by anyone before publishing, because there is no validator in anybody's build for this block, and 25 of them shipped a string that a five line check would have caught.
-
product, product.groupExtended vocabulary Four pages. Real types in Facebook's extended set, not in the thirteen the specification lists as global. -
Homepage, homepage, Page, page, site, webPage description Eight pages declared a word for what the page is rather than a type of object. -
Page d'accueilPage description Two pages. The same mistake in French, which suggests a localised template string reached the field. -
illinois-home-page, BECU:homeInternal identifier A page identifier and a namespaced content key, both leaked straight out of a content management system. -
Web Area Homepage, Education, Articles, CompanySection label Four pages bound og:type to a navigation section or a content category. -
object, ObjectCopied from the prose Two pages declared the word the specification uses to describe what og:type applies to. -
webisteMisspelling One page. website with two characters transposed, which no build step caught. -
hospiTruncated string One page, hospitalitaliano.org.ar, which also shipped the unfilled og:url placeholder.
27 pages wrote Open Graph into the attribute the specification does not use
Every example in the Open Graph protocol is written with a property attribute: property="og:title" carrying a content value. That choice is not arbitrary. The HTML Standard's meta element defines exactly five content attributes, name, http-equiv, content, charset and media, and property is not among them; it arrives from RDFa, which is the vocabulary mechanism Open Graph is built on. A parser implementing the specification looks for the property attribute, and a page that writes name="og:title" has published a meta element that the HTML Standard does define and that the Open Graph protocol does not describe.
27 of the 1,135 pages did that. On 15 of them it was duplication or a stray tag, with a correct property-based block also present, and nothing is lost. On the other 12 there was no property-based og:title anywhere in the head at all. Two of those 12 shipped a complete block in the wrong attribute: dropbox.com wrote og:title, og:description, og:url, og:type, og:site_name, og:image, og:image:width and og:image:height all as name attributes, and cedars-sinai.org wrote og:locale, og:type, og:site_name, og:url, og:title, og:description and og:image the same way. nationwidechildrens.org wrote four, backmarket.fr three, northwesternmutual.com, langhamhotels.com and ryanmulligan.dev two each, and finnair.com, endclothing.com, mandarinoriental.com, inquirer.com and spelman.edu one each.
This is the most consequential of the four findings here and also the hardest to see, which is why it is worth being careful about what it does and does not mean. A developer opening the page source, or the elements panel, sees a full Open Graph block and has no reason to look at which attribute carries it. A consumer that scans for any meta tag whose name or property starts with og: will read these pages perfectly well, and plenty of software is written that way. A consumer implementing the specification as written will find nothing. Which behaviour a given AI answer engine implements is not published by anyone, so the measurable fact is the divergence rather than its cost: 12 pages in this corpus are, to a strict reader, Open Graph free while presenting as fully marked up.
It is the same class of failure as declaring markup in a form a parser does not accept, which this corpus has now produced in three separate shapes: 22 of 385 home pages carrying microdata and one carrying RDFa where the consuming documentation asks for JSON-LD, 429 of 615 pages naming no language in their JSON-LD, and now a block written in the wrong attribute. The practical version of this for a reader who wants to act on it, rather than a measurement of other people's sites, is in the guide to getting cited in Google AI Overviews.
What the specification asks for
- meta property="og:title"
- meta property="og:type"
- meta property="og:image"
- meta property="og:url"
- 578 of 948 pages carried all four this way
dropbox.com and cedars-sinai.org, as returned
- meta name="og:title"
- meta name="og:type"
- meta name="og:image"
- meta name="og:url"
- 12 pages carried no property-based og:title
What this measurement does not show
Six limits, and they bound every figure above. The first is the one that matters most: no crawler was observed. This run made one HTTPS request per hostname from one network location and read the bytes that came back. It holds no server logs, it inspected no retrieval pipeline, and it therefore supports no claim that GPTBot, ClaudeBot, PerplexityBot or any other named bot read, ignored or was confused by any of these tags. Google's title link documentation is the single published statement that any of this block is consumed, and it covers og:title alone.
The second is that a defect here is not a measured cost. A private address in an og:url is wrong on its face and nobody has published what a consumer does about it. The honest claim is that the page states something untrue about itself, not that the statement was acted on, and the distinction is the same one drawn in the research index for every external study cited there.
The third is sampling. The 1,419 hostnames are an editorial frame assembled for platform and industry coverage, not a random draw of the web, so every rate here describes these hostnames and nothing wider. The strata are also uneven in what they contributed: 284 hosts did not answer 200 with HTML, 220 of them answering 403, and defended sites are under-represented in what remains for exactly that reason.
The fourth is that this was one request per host, unlike the two-pass design used when a finding turned on a single page's behaviour. The fourteen og:url defects and the twelve name-attribute pages were each confirmed by a second independent request while writing this post, but the aggregate counts rest on one read, and a site serving different markup to a second request would not have been caught.
The fifth is that no JavaScript ran, which for this block is less of a limit than it is elsewhere but is not nothing: an earlier run found 27 of 404 home pages with no structured data in the HTML gained some when a browser executed the page, and a head tag injected by script would be missed here the same way. That can only make the missing counts too high.
The sixth is scope. Home pages only, one page per site, and og:image URLs were recorded but never fetched, so nothing here says whether those images resolve. Anyone who wants the same read of their own pages can get it from what a crawler sees on a URL, and the neighbouring question of which external identifiers a page claims was measured separately as 85 of 317 organizations naming a reference entry in sameAs.
Sample Illustrative, not a measurement of any real site.
Flow: Read the head, no JavaScript to Is it in a property attribute?; Is it in a property attribute? (yes) to All four required present?; Is it in a property attribute? (no, 12 pages) to Block states something untrue; All four required present? (yes, 578) to Is og:type a defined type?; All four required present? (no, 370) to Block states something untrue; Is og:type a defined type? (yes, 751) to Does og:url resolve to this page?; Is og:type a defined type? (no, 25) to Block states something untrue; Does og:url resolve to this page? (yes, 753) to Block is well formed; Does og:url resolve to this page? (no, 14) to Block states something untrue.
Lantad
Published .
The Open Graph block is the oldest machine-readable layer on a normal web page and the one nobody looks at twice. It was built in 2010 so a social network could draw a card, it outlived the reason it was built, and it now sits in the head of most commercial sites as a second, parallel description of what the page is: a name, a type, an image and a canonical address, written in tags that no browser renders and no visitor ever sees. That last part is what makes it worth measuring. A title in the browser tab is checked by everybody who opens the site. An og:url is checked by nobody, because nothing on the page changes when it is wrong.
Common questions
Do AI crawlers read Open Graph tags?
Only one vendor publishes a commitment. Google's title link documentation, carrying Last updated 2025-12-10 UTC, names content in og:title meta tags as one of nine sources it uses to choose a title link. No operator documentation behind the crawler tokens this scanner evaluates names og:type, og:url or og:site_name, so for every other engine and every other property the position is undocumented rather than known. This run fetched pages and observed no crawler, so it adds nothing to that question beyond what is in the bytes.
Which Open Graph tags are required?
The Open Graph protocol specification names four required properties for every page: og:title, og:type, og:image and og:url. Of the 948 home pages in this corpus carrying at least one og: property on 29 September 2026, 578 carried all four and 370 did not. og:url was the one most often missing, absent on 181 of the 948.
Should Open Graph use the property attribute or the name attribute?
The specification uses property, and every example in it is written that way. The HTML Standard's meta element defines five content attributes, name, http-equiv, content, charset and media, and property is not among them; it comes from RDFa. 27 of the 1,135 pages read here used a name attribute for og: names, and on 12 of those there was no property-based og:title in the head at all, so a parser implementing the specification as written finds no Open Graph on those pages.
What happens if og:url points at the wrong address?
Nobody publishes what a consumer does about it, so the honest answer is that the page states a false permanent identifier for itself and the consequence is undocumented. On 14 of the 767 pages declaring an og:url in this corpus, the value did not resolve to the page that served it, including a private RFC 1918 address, a vendor-internal hostname, a hosting provider's staging hostname and an unfilled template placeholder reading https://www.your-domain.com/your-page.html.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.