BlogFindings

JavaScript structured data: 27 of 404 home pages with none in the HTML gained some when a browser ran the page

Lantad requested robots.txt from all 1,419 hostnames in this repository's committed corpus on 29 September 2026, then read every permitted home page twice: once as raw HTML with no JavaScript executed, and once out of the DOM of a real browser sent the same crawler user agent. 969 hosts produced a clean pair. 404 of those carried no JSON-LD block in the HTML, and 27 of the 404 carried one after the browser had run the page.

17 min read Lantad

This run measured how often the distinction actually bites. Lantad requested /robots.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026 and evaluated each file for its own crawler token at the site root before asking for anything else. Every permitted home page was then fetched twice. The first request was a plain HTTPS fetch with no JavaScript executed, which is what a non-rendering crawler reads. The second loaded the same URL in a headless Chromium sent the same crawler user agent, waited for the load event and a fixed settle period, and took the JSON-LD out of the live DOM. Nothing differed between the two reads except whether scripts ran, which is the only way to attribute a difference to JavaScript rather than to the server.

In short

  • JavaScript structured data changed the JSON-LD a crawler could read on 52 of the 969 home pages Lantad read twice on 29 September 2026, and left the block count identical on 907 of them.
  • 404 of those 969 home pages carried no JSON-LD block in the raw HTML. 27 of the 404 carried at least one after a browser ran the page and 374 returned no block in either read, so on this corpus an absent block is far more often absent than hidden.
  • Google's guidance on generating structured data with JavaScript, carrying Last updated 2025-12-10 UTC, states that Google Search can understand and process structured data that's available in the DOM when it renders the page.
  • The 27 pages whose markup existed only after JavaScript were mostly large institutions rather than tag manager installs: 6 sat in the finance stratum, 6 in healthcare and 4 in SaaS, and only 2 of the 27 loaded Google Tag Manager or gtag in their HTML at all.
  • 25 further home pages added blocks on top of what the server already sent, and those additions cluster on ratings: across the 52 sites that changed, AggregateRating appeared only after JavaScript on 8, Review on 5 and FAQPage on 5.
StageCountWhat happened
Hostnames asked1,419The committed corpus, an editorial frame rather than a random draw
robots.txt never returned a status59No home page was requested for these
Disallowed LantadBot at the root5212 by an explicit rule, 40 because robots.txt answered 503, which this parser treats as disallow
Answered HTTP 202, body not parsed7A gap in this run's harness, so these were left out rather than measured
Home pages requested1,301Every host whose robots.txt allowed the root
Did not return both reads257216 answered 403, 7 answered 429, 4 answered 404, 13 renders failed twice and 17 returned another status or no body
Pairs obtained1,044A raw fetch and a render for the same URL
Excluded as unreliable7524 by a challenge page title, 43 by a degraded render, 8 by a firewall interstitial
Clean comparisons969The denominator for every figure in this post
One GET of https://<host>/robots.txt, then two reads of https://<host>/: a plain HTTPS fetch with no JavaScript, and a headless Chromium load. Both were sent as LantadBot/1.0 (+https://lantad.co/bot) with redirects followed, a twenty five second timeout on the fetch and a twenty second load budget plus a three second settle on the render, from one network location. robots.txt was parsed with parseRobotsTxt and evaluated for LantadBot at the site root with evaluateRobots, both imported directly from core/src/robots.ts. Measured by Lantad on 29 September 2026 across the 1,419 hostnames in worker/seeds/corpus-seeds-platform.json and worker/seeds/corpus-seeds-industry.json.

What does Google say about JavaScript structured data, and what do AI crawler operators say?

One of the two audiences for this markup has published a clear position. Google's page on generating structured data with JavaScript, carrying Last updated 2025-12-10 UTC, describes using JavaScript "to either generate all of your structured data or add more information to the server-side rendered structured data" and then states: "Either way, Google Search can understand and process structured data that's available in the DOM when it renders the page." Google's introduction to structured data, carrying the same date, says the same thing from the other side: "Google can read JSON-LD data when it is dynamically injected into the page's contents, such as by JavaScript code or embedded widgets in your content management system."

Two things in that are worth holding onto. The first is that Google names exactly two patterns, generating all of it and adding to what the server sent, and the counts below are organised around that split because the documentation already drew it. The second is that the promise is conditional on rendering. It is a statement about a crawler that runs a browser, and it says nothing about one that does not.

The other audience has published almost nothing. When Lantad fetched the nine vendor documentation pages behind the fifteen crawler tokens in its registry on 5 September 2026, 2 of the 9 stated whether their crawler executes JavaScript and 7 did not. Common Crawl's FAQ ruled it out in one sentence and Apple's Applebot page ruled it in. The rest are silent, which means a site whose markup exists only after a script has run is not doing something documented as broken. It is relying on behaviour that most of the operators it cares about have never described. That is a different and more awkward position than either working or failing, and it is the reason this measurement counts occurrences rather than predicting outcomes.

Google, in writing

  • Either way, Google Search can
  • understand and process structured data
  • that's available in the DOM when it
  • renders the page.
  • Google can read JSON-LD data when it is
  • dynamically injected into the page's
  • contents.
  • Says: rendering is done, and read.

AI crawler operators

  • 9 vendor documentation pages behind
  • 15 crawler tokens, read 5 Sept 2026.
  • 2 state a position on JavaScript.
  • 7 state nothing either way.
  • Common Crawl rules it out.
  • Applebot rules it in.
  • Says: mostly nothing.
The published positions, quoted exactly as read at source. The two Google pages were read on 29 September 2026 and both carry Last updated 2025-12-10 UTC. The vendor documentation count is Lantad's own, measured on 5 September 2026.

How much structured data was in the HTML before any JavaScript ran

The raw fetch comes first because it is the weaker reader, and what it sees is the floor. Of the 969 home pages with a clean pair, 565 carried at least one script element of type application/ld+json in the bytes the server sent and 404 carried none. Those 565 pages held 1,072 blocks between them. 15 pages served a block that did not parse as JSON, which is the same defect class this corpus counted when it found 117 of 612 home pages with JSON-LD carrying a defect.

That 404 is the number the rest of the post turns on, so it is worth being careful about what it is. It is not a count of sites with no structured data, because some of them describe themselves in microdata or RDFa instead, and it is not a count of sites doing something wrong. It is the population for whom the question has any force: if a JSON-LD block is missing from the HTML, either it does not exist, or it exists and arrives later. The split between those two is measurable and nobody appears to have measured it at this scale.

The raw read also turned out to be stable, which matters more than it sounds. Every one of the 969 hosts was fetched a second time later the same day, and 968 of them returned the identical JSON-LD block count. A census that moves between two reads hours apart cannot support a claim about a difference of 27 sites, so that check was run before any of the figures below were written rather than after. The earlier corpus run that found 141 of 382 home pages carrying no structured data in the raw HTML measured the same property on a smaller frame and produced a similar proportion.

  • At least one JSON-LD block in the raw HTML 565 of 969 Holding 1,072 blocks between them
  • No JSON-LD block in the raw HTML 404 of 969 The population where JavaScript could be the only source
  • Loaded Google Tag Manager or gtag 209 of 969 75 of these were pages carrying no JSON-LD block
  • Served a JSON-LD block that did not parse 15 of 969 Markup published in a form no consumer can read
  • Identical block count on a second fetch 968 of 969 The stability check the difference figures rest on
What the 969 home pages carried, counted from the raw bytes of one request per host with no JavaScript executed, on 29 September 2026. Each figure is a count of what the markup declares, not a measure of what any engine does with it.

27 of the 404 pages with no markup in the HTML gained some when a browser ran them

This is Google's first pattern, generating all of your structured data in JavaScript, and it is rarer than its reputation. 30 of the 404 pages carrying no JSON-LD in the HTML produced at least one parseable block in the rendered DOM. Every one of those 30 was then measured again from scratch, and 27 reproduced. The three that did not, byword.ai, stuff.co.nz and gymshark.com, returned no block in either read the second time, so they are reported here as unreproduced rather than counted. The remaining 374 of the 404 returned no block in either read, which is the plain finding: on this corpus, markup missing from the HTML is usually just missing.

The 27 are not who the common advice implies. The story usually told about JavaScript-injected schema is a tag manager story, and the tag manager is almost absent here: 2 of the 27 loaded Google Tag Manager or gtag in their HTML at all. What the list actually holds is large institutions whose home pages assemble themselves in the client. 6 are in the finance stratum, among them visa.com, allstate.com, progressive.com, scotiabank.com, northwesternmutual.com and bradesco.com.br. 6 are in healthcare, including massgeneral.org, dana-farber.org, amsterdamumc.nl, nice.org.uk and rcgp.org.uk. 4 are SaaS, including xero.com, 1password.com and cockroachlabs.com. europa.eu is there too, publishing a GovernmentOrganization and a breadcrumb trail that exist only once scripts have run.

What those pages put in JavaScript is mostly identity rather than decoration. Organization or Corporation nodes, a postal address, a contact point, a WebSite with a search action: the block that tells a machine who the publisher is. That is the same property this scanner reads for entity confidence, and it is the part of a page a non-rendering crawler can least afford to miss, because there is no second place on a home page where the publisher's identity is stated in a machine-readable form. Whether any given crawler misses it is not something this run can say, and the section below on limits is explicit about that.

The comparison each of the 969 hosts went through, and how many landed in each outcome. Measured by Lantad on 29 September 2026. The two branches partition their inputs exactly: 25 plus 533 plus 7 is 565, and 27 plus 3 plus 374 is 404. The 75 unreliable pairs were removed before the 969 was fixed, not after.

25 sites added markup on top of what the server sent, and the additions were ratings

Google's second pattern, adding more information to the server-side rendered structured data, appeared on 25 of the 565 pages that already carried a block. All 25 reproduced on a second measurement. These are the smaller sites, and here the tag manager really is involved: 10 of the 25 loaded Google Tag Manager or gtag, against 2 of the 27 in the previous section. The direction of the difference is the interesting part. Enterprise sites hide their identity markup in the client by accident of architecture. Small business sites bolt extra markup on through a widget, on purpose.

What gets bolted on is consistent enough to name. Across the 52 sites whose markup changed, AggregateRating appeared only after JavaScript on 8, Review on 5, Rating on 4 and Product and Brand on 4 each. That cluster is a reviews widget writing its own block, and it shows up on jimmyjoesplumbing.com, bergenpronotary.com, prenvalleybuilders.com, westashevillefamilydentistry.com and grandstreetdental.com, every one of them a local trade or practice on Wix, Squarespace or WordPress. The same shape this corpus met when it counted 33 of 615 home pages declaring a rating is partly, on these sites, not in the HTML at all.

Two other clusters recur. FAQPage with its Question and Answer nodes arrived only after JavaScript on 5 sites, which is the markup behind the pattern counted when this corpus found 298 of 1,277 declared answers were not on the page. VideoObject arrived late on 4, including collibra.com, workday.com and spellbook.legal, the last of which went from 1 block to 9 when its player scripts ran, and that is the same under-declaration counted when only 21 of the 270 home pages showing a video declared one. zendesk.com went from 1 block to 7, adding Person and Review nodes. A crawler that does not render sees the company and not the ratings, the questions or the video, which is a narrower loss than losing everything and a more commercially specific one.

HostStratumBlocks, raw to renderedGTMTypes only in the DOM
spellbook.legalwebflow1 to 9noClip, VideoObject
zendesk.comsaas1 to 7noPerson, Review
collibra.comsaas3 to 8noVideoObject
workday.comsaas1 to 4noVideoObject
hardawayvet.commedia-local1 to 4noAggregateRating, FAQPage, LocalBusiness, Person, Question
westashevillefamilydentistry.comwordpress-smb3 to 4yesAggregateRating, Brand, Rating, Review
jimmyjoesplumbing.comwix-squarespace3 to 4yesAggregateRating, Brand, Person, Product, Rating, Review
cuure.combubble-nocode2 to 4no
predictap.comsaas-marketing1 to 3yesAggregateRating
oliverbonas.comecommerce1 to 3noContactPoint, Organization, PostalAddress, SearchAction, WebSite
bergenpronotary.comwix-squarespace1 to 2yesAggregateRating, Brand, Person, Product, Rating, Review
berkeleyscanner.commedia-local1 to 2yesNewsMediaOrganization, Person, PostalAddress, City
hometownplumbingtn.comwordpress-smb1 to 2yesLocalBusiness, Offer, OfferCatalog, Organization, Place, Service
trueform.agencyframer1 to 2noAnswer, FAQPage, Question
trinityschoolofmedicine.orgwebflow1 to 2noAnswer, FAQPage, Question
Plus 10 moremixed1 to 2 or 2 to 3mixedgrandstreetdental.com, prenvalleybuilders.com, whipsaw.com, thewireguyelectric.com, quartr.com, scaleway.com, eltiempo.com, eurostar.com, france.fr, ally.com
Every site measured twice where a browser produced more JSON-LD blocks than the server sent, with the block counts from the confirming second measurement. Measured by Lantad on 29 September 2026. New types are @type values present in the rendered DOM and absent from the raw HTML; a blank means the extra block repeated types already present.

Which stacks put their markup where, across eighteen strata

The corpus is stratified twice over, by the platform a site is built on and by the industry it trades in, and both cuts say something the aggregate hides. The platform cut is the cleaner story about causation, because the platform is what decides whether prose and markup reach a crawler at all. 50 of the 51 Wix and Squarespace pages and 26 of the 27 WordPress pages carried a JSON-LD block in the HTML: those platforms emit it server side and the question barely arises. At the other end, 23 of the 30 static documentation sites, 28 of the 39 no code sites and 22 of the 42 Webflow sites carried none, and in most of those cases none arrived after rendering either.

The stratum where client-side assembly is most expected turned out to be the quietest. 14 of the 40 single page application startups carried no block in the HTML, and not one of the 40 gained a block from rendering, so the architecture that runs the most JavaScript was not the architecture hiding markup behind it. That is consistent with what this corpus found measuring prose rather than markup, where JavaScript supplied 7.6 percent of the words and all of it on 11 of 271 pages, and with the run where rendering added no new crawl paths. A page can be built entirely in the browser and still ship no schema in either view.

The industry cut is where the 27 concentrate. Finance and healthcare supplied 12 of them between 168 pages, against 0 from 51 news sites, and news is the industry stratum least likely to be missing markup in the HTML, with 5 of 51 carrying none. Regulated industries running large content management systems behind client-side front ends are the profile, which is the same population whose organization node is missing or unnamed in earlier runs. If you run one of these stacks, the per-stack notes for Next.js and React cover where the markup goes, and what GPTBot sees will show you your own page as a non-rendering fetch returns it.

StratumClean pairsNo JSON-LD in HTMLGainedAdded
industry/saas1142744
industry/education916510
industry/finance853361
industry/healthcare834560
industry/government825410
industry/travel663232
industry/ecommerce511721
industry/news51501
platform/wix-squarespace51103
platform/webflow422214
platform/spa-startups401400
platform/bubble-nocode392811
platform/saas-marketing341121
platform/shopify-dtc31900
platform/static-docs302300
platform/wordpress-smb27103
platform/framer261602
platform/media-local26102
The eighteen strata of the committed corpus, restricted to the 969 hosts with a clean pair on 29 September 2026. Gained is Google's first pattern, a page with no JSON-LD in the HTML that produced some after rendering; added is its second, a page that produced more blocks than the server sent. Both columns count only sites that reproduced on a second measurement.

What this run did not measure, and what it threw away

The largest limit is the one the figures cannot cross. No crawler was observed reading any of these pages, and no server logs were held, so nothing here supports a claim that GPTBot, ClaudeBot, PerplexityBot or any other named bot missed a block. The headless Chromium used for the second read stands in for a rendering client and is not any vendor's crawler. What this run establishes is where the markup is, which is a fact about the sites. What an engine does about it is a separate question that the methodology page is explicit about not answering, and the honest form of the finding is the conditional one: a client that does not render sees the HTML column and nothing else.

The second limit is that the exclusions were not random. 75 of the 1,044 pairs were dropped: 24 whose title matched a challenge signature, 13 of them the string "One moment, please...", 43 whose render produced a degraded document, and 8 whose raw body carried a web application firewall interstitial with fewer than 50 words of text, among them pennmedicine.org, zurich.com and gtbank.com at 212 bytes each. Those pages were served a bot wall rather than a page, and counting a bot wall as a site with no structured data would have manufactured the finding instead of measuring it. An earlier pass of this analysis did exactly that, and the correction removed 10 sites that had looked like the headline. The exclusions concentrate in strata that sit behind firewalls, so the remaining 969 under-represents heavily defended sites, which is a bias toward sites that let a crawler in.

Three narrower limits are worth stating. The render was given a fixed budget of the load event plus three seconds, so markup injected later than that was not seen, which can only make the 27 and the 25 too low rather than too high. 7 hostnames answered robots.txt with HTTP 202 and this run did not parse those bodies, so they were left out rather than measured, and that is a gap in the harness rather than a property of those sites. And 7 of the 969 produced fewer blocks after rendering than before, which would be the interesting inverse finding if it could be separated from a script failing to load behind this network; it cannot be, from one location with one attempt, so no claim is made from it. Everything above describes these 1,419 hostnames on one day, and the corpus is an editorial sampling frame assembled for coverage rather than a random draw of the web.

  • Where the markup is Measured Two reads of the same URL differing only in whether scripts ran, on 969 hosts, with the raw read reproduced on 968 of them.
  • Which crawler missed it Not measured No crawler was observed and no server logs were held. The browser used here is not any vendor's bot.
  • Whether it costs citations Not measured Nothing in this run links a block's location to an appearance in any answer engine.
  • Late injected markup Undercounted The render budget was the load event plus three seconds, so anything slower was missed. This biases the counts down.
  • Heavily defended sites Under-represented 75 pairs were excluded as challenge pages or degraded renders, and those cluster behind firewalls rather than at random.
  • Markup lost to JavaScript Withheld 7 pages produced fewer blocks after rendering, indistinguishable from a script blocked on this network, so no figure is claimed.
What this run can and cannot support, stated before the conclusions rather than after them. Measured by Lantad on 29 September 2026.

Written by

Lantad

Published .

A page can carry structured data in two places. It can sit in the bytes the server sends, where anything that reads HTML will find it, or it can be written into the document by JavaScript after the page loads, where only a client that executes scripts will ever see it. JavaScript structured data is the second case, and it matters here because the two populations of reader are not the same. Google says in writing that it renders pages and reads what rendering produces. Most AI crawler operators say nothing on the subject in either direction.

Common questions

Does JavaScript structured data work for AI search?

Google states that it does for Google Search: its guidance on generating structured data with JavaScript, carrying Last updated 2025-12-10 UTC, says Google Search can understand and process structured data that's available in the DOM when it renders the page. For AI crawlers the answer is undocumented rather than known. Of the 9 vendor documentation pages behind the 15 crawler tokens Lantad tracks, read on 5 September 2026, 2 stated whether their crawler executes JavaScript and 7 stated nothing either way.

Is my schema invisible because a script injects it?

Usually not, on the evidence of this run. Of 969 home pages Lantad read twice on 29 September 2026, 907 carried the identical number of JSON-LD blocks before and after a browser ran them. The pages where JavaScript was the only source of any block numbered 27, out of 404 that carried none in the HTML. The more common case by a wide margin, 374 of those 404, is a page with no structured data in either read.

Which schema types are most often added by JavaScript?

Identity and ratings, in that order. Across the 52 sites whose block count changed on 29 September 2026, Organization appeared only in the rendered DOM on 21, WebSite on 15, PostalAddress on 14 and ContactPoint on 13. The ratings cluster follows: AggregateRating on 8, Review on 5, Rating on 4, and Product and Brand on 4 each. FAQPage appeared late on 5 sites and VideoObject on 4.

How do I check whether my own structured data is in the HTML?

Request your own page without a browser and look for script elements of type application/ld+json in the bytes that come back, then load the same URL in a browser and count them again. A difference is markup that only a rendering client will ever read. Lantad's what GPTBot sees tool performs the non-rendering half of that comparison on a single URL, and its scan reports the same split across a site.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.