BlogFindings
JavaScript structured data: 27 of 404 home pages with none in the HTML gained some when a browser ran the page
Lantad requested robots.txt from all 1,419 hostnames in this repository's committed corpus on 29 September 2026, then read every permitted home page twice: once as raw HTML with no JavaScript executed, and once out of the DOM of a real browser sent the same crawler user agent. 969 hosts produced a clean pair. 404 of those carried no JSON-LD block in the HTML, and 27 of the 404 carried one after the browser had run the page.
This run measured how often the distinction actually bites. Lantad requested /robots.txt from all 1,419 hostnames in this repository's two committed corpus seed files on 29 September 2026 and evaluated each file for its own crawler token at the site root before asking for anything else. Every permitted home page was then fetched twice. The first request was a plain HTTPS fetch with no JavaScript executed, which is what a non-rendering crawler reads. The second loaded the same URL in a headless Chromium sent the same crawler user agent, waited for the load event and a fixed settle period, and took the JSON-LD out of the live DOM. Nothing differed between the two reads except whether scripts ran, which is the only way to attribute a difference to JavaScript rather than to the server.
In short
- JavaScript structured data changed the JSON-LD a crawler could read on 52 of the 969 home pages Lantad read twice on 29 September 2026, and left the block count identical on 907 of them.
- 404 of those 969 home pages carried no JSON-LD block in the raw HTML. 27 of the 404 carried at least one after a browser ran the page and 374 returned no block in either read, so on this corpus an absent block is far more often absent than hidden.
- Google's guidance on generating structured data with JavaScript, carrying Last updated 2025-12-10 UTC, states that Google Search can understand and process structured data that's available in the DOM when it renders the page.
- The 27 pages whose markup existed only after JavaScript were mostly large institutions rather than tag manager installs: 6 sat in the finance stratum, 6 in healthcare and 4 in SaaS, and only 2 of the 27 loaded Google Tag Manager or gtag in their HTML at all.
- 25 further home pages added blocks on top of what the server already sent, and those additions cluster on ratings: across the 52 sites that changed, AggregateRating appeared only after JavaScript on 8, Review on 5 and FAQPage on 5.
| Stage | Count | What happened |
|---|---|---|
| Hostnames asked | 1,419 | The committed corpus, an editorial frame rather than a random draw |
| robots.txt never returned a status | 59 | No home page was requested for these |
| Disallowed LantadBot at the root | 52 | 12 by an explicit rule, 40 because robots.txt answered 503, which this parser treats as disallow |
| Answered HTTP 202, body not parsed | 7 | A gap in this run's harness, so these were left out rather than measured |
| Home pages requested | 1,301 | Every host whose robots.txt allowed the root |
| Did not return both reads | 257 | 216 answered 403, 7 answered 429, 4 answered 404, 13 renders failed twice and 17 returned another status or no body |
| Pairs obtained | 1,044 | A raw fetch and a render for the same URL |
| Excluded as unreliable | 75 | 24 by a challenge page title, 43 by a degraded render, 8 by a firewall interstitial |
| Clean comparisons | 969 | The denominator for every figure in this post |
What does Google say about JavaScript structured data, and what do AI crawler operators say?
One of the two audiences for this markup has published a clear position. Google's page on generating structured data with JavaScript, carrying Last updated 2025-12-10 UTC, describes using JavaScript "to either generate all of your structured data or add more information to the server-side rendered structured data" and then states: "Either way, Google Search can understand and process structured data that's available in the DOM when it renders the page." Google's introduction to structured data, carrying the same date, says the same thing from the other side: "Google can read JSON-LD data when it is dynamically injected into the page's contents, such as by JavaScript code or embedded widgets in your content management system."
Two things in that are worth holding onto. The first is that Google names exactly two patterns, generating all of it and adding to what the server sent, and the counts below are organised around that split because the documentation already drew it. The second is that the promise is conditional on rendering. It is a statement about a crawler that runs a browser, and it says nothing about one that does not.
The other audience has published almost nothing. When Lantad fetched the nine vendor documentation pages behind the fifteen crawler tokens in its registry on 5 September 2026, 2 of the 9 stated whether their crawler executes JavaScript and 7 did not. Common Crawl's FAQ ruled it out in one sentence and Apple's Applebot page ruled it in. The rest are silent, which means a site whose markup exists only after a script has run is not doing something documented as broken. It is relying on behaviour that most of the operators it cares about have never described. That is a different and more awkward position than either working or failing, and it is the reason this measurement counts occurrences rather than predicting outcomes.
Google, in writing
- Either way, Google Search can
- understand and process structured data
- that's available in the DOM when it
- renders the page.
- Google can read JSON-LD data when it is
- dynamically injected into the page's
- contents.
- Says: rendering is done, and read.
AI crawler operators
- 9 vendor documentation pages behind
- 15 crawler tokens, read 5 Sept 2026.
- 2 state a position on JavaScript.
- 7 state nothing either way.
- Common Crawl rules it out.
- Applebot rules it in.
- Says: mostly nothing.
How much structured data was in the HTML before any JavaScript ran
The raw fetch comes first because it is the weaker reader, and what it sees is the floor. Of the 969 home pages with a clean pair, 565 carried at least one script element of type application/ld+json in the bytes the server sent and 404 carried none. Those 565 pages held 1,072 blocks between them. 15 pages served a block that did not parse as JSON, which is the same defect class this corpus counted when it found 117 of 612 home pages with JSON-LD carrying a defect.
That 404 is the number the rest of the post turns on, so it is worth being careful about what it is. It is not a count of sites with no structured data, because some of them describe themselves in microdata or RDFa instead, and it is not a count of sites doing something wrong. It is the population for whom the question has any force: if a JSON-LD block is missing from the HTML, either it does not exist, or it exists and arrives later. The split between those two is measurable and nobody appears to have measured it at this scale.
The raw read also turned out to be stable, which matters more than it sounds. Every one of the 969 hosts was fetched a second time later the same day, and 968 of them returned the identical JSON-LD block count. A census that moves between two reads hours apart cannot support a claim about a difference of 27 sites, so that check was run before any of the figures below were written rather than after. The earlier corpus run that found 141 of 382 home pages carrying no structured data in the raw HTML measured the same property on a smaller frame and produced a similar proportion.
27 of the 404 pages with no markup in the HTML gained some when a browser ran them
This is Google's first pattern, generating all of your structured data in JavaScript, and it is rarer than its reputation. 30 of the 404 pages carrying no JSON-LD in the HTML produced at least one parseable block in the rendered DOM. Every one of those 30 was then measured again from scratch, and 27 reproduced. The three that did not, byword.ai, stuff.co.nz and gymshark.com, returned no block in either read the second time, so they are reported here as unreproduced rather than counted. The remaining 374 of the 404 returned no block in either read, which is the plain finding: on this corpus, markup missing from the HTML is usually just missing.
The 27 are not who the common advice implies. The story usually told about JavaScript-injected schema is a tag manager story, and the tag manager is almost absent here: 2 of the 27 loaded Google Tag Manager or gtag in their HTML at all. What the list actually holds is large institutions whose home pages assemble themselves in the client. 6 are in the finance stratum, among them visa.com, allstate.com, progressive.com, scotiabank.com, northwesternmutual.com and bradesco.com.br. 6 are in healthcare, including massgeneral.org, dana-farber.org, amsterdamumc.nl, nice.org.uk and rcgp.org.uk. 4 are SaaS, including xero.com, 1password.com and cockroachlabs.com. europa.eu is there too, publishing a GovernmentOrganization and a breadcrumb trail that exist only once scripts have run.
What those pages put in JavaScript is mostly identity rather than decoration. Organization or Corporation nodes, a postal address, a contact point, a WebSite with a search action: the block that tells a machine who the publisher is. That is the same property this scanner reads for entity confidence, and it is the part of a page a non-rendering crawler can least afford to miss, because there is no second place on a home page where the publisher's identity is stated in a machine-readable form. Whether any given crawler misses it is not something this run can say, and the section below on limits is explicit about that.
Flow: Raw fetch, no JavaScript to JSON-LD in HTML: 565; Raw fetch, no JavaScript to No JSON-LD in HTML: 404; JSON-LD in HTML: 565 (on top) to More after JS: 25; JSON-LD in HTML: 565 to Identical: 533; JSON-LD in HTML: 565 to Fewer after JS: 7; No JSON-LD in HTML: 404 (only source) to Only after JS: 27; No JSON-LD in HTML: 404 to Not reproduced: 3; No JSON-LD in HTML: 404 to None in either read: 374.
25 sites added markup on top of what the server sent, and the additions were ratings
Google's second pattern, adding more information to the server-side rendered structured data, appeared on 25 of the 565 pages that already carried a block. All 25 reproduced on a second measurement. These are the smaller sites, and here the tag manager really is involved: 10 of the 25 loaded Google Tag Manager or gtag, against 2 of the 27 in the previous section. The direction of the difference is the interesting part. Enterprise sites hide their identity markup in the client by accident of architecture. Small business sites bolt extra markup on through a widget, on purpose.
What gets bolted on is consistent enough to name. Across the 52 sites whose markup changed, AggregateRating appeared only after JavaScript on 8, Review on 5, Rating on 4 and Product and Brand on 4 each. That cluster is a reviews widget writing its own block, and it shows up on jimmyjoesplumbing.com, bergenpronotary.com, prenvalleybuilders.com, westashevillefamilydentistry.com and grandstreetdental.com, every one of them a local trade or practice on Wix, Squarespace or WordPress. The same shape this corpus met when it counted 33 of 615 home pages declaring a rating is partly, on these sites, not in the HTML at all.
Two other clusters recur. FAQPage with its Question and Answer nodes arrived only after JavaScript on 5 sites, which is the markup behind the pattern counted when this corpus found 298 of 1,277 declared answers were not on the page. VideoObject arrived late on 4, including collibra.com, workday.com and spellbook.legal, the last of which went from 1 block to 9 when its player scripts ran, and that is the same under-declaration counted when only 21 of the 270 home pages showing a video declared one. zendesk.com went from 1 block to 7, adding Person and Review nodes. A crawler that does not render sees the company and not the ratings, the questions or the video, which is a narrower loss than losing everything and a more commercially specific one.
| Host | Stratum | Blocks, raw to rendered | GTM | Types only in the DOM |
|---|---|---|---|---|
| spellbook.legal | webflow | 1 to 9 | no | Clip, VideoObject |
| zendesk.com | saas | 1 to 7 | no | Person, Review |
| collibra.com | saas | 3 to 8 | no | VideoObject |
| workday.com | saas | 1 to 4 | no | VideoObject |
| hardawayvet.com | media-local | 1 to 4 | no | AggregateRating, FAQPage, LocalBusiness, Person, Question |
| westashevillefamilydentistry.com | wordpress-smb | 3 to 4 | yes | AggregateRating, Brand, Rating, Review |
| jimmyjoesplumbing.com | wix-squarespace | 3 to 4 | yes | AggregateRating, Brand, Person, Product, Rating, Review |
| cuure.com | bubble-nocode | 2 to 4 | no | |
| predictap.com | saas-marketing | 1 to 3 | yes | AggregateRating |
| oliverbonas.com | ecommerce | 1 to 3 | no | ContactPoint, Organization, PostalAddress, SearchAction, WebSite |
| bergenpronotary.com | wix-squarespace | 1 to 2 | yes | AggregateRating, Brand, Person, Product, Rating, Review |
| berkeleyscanner.com | media-local | 1 to 2 | yes | NewsMediaOrganization, Person, PostalAddress, City |
| hometownplumbingtn.com | wordpress-smb | 1 to 2 | yes | LocalBusiness, Offer, OfferCatalog, Organization, Place, Service |
| trueform.agency | framer | 1 to 2 | no | Answer, FAQPage, Question |
| trinityschoolofmedicine.org | webflow | 1 to 2 | no | Answer, FAQPage, Question |
| Plus 10 more | mixed | 1 to 2 or 2 to 3 | mixed | grandstreetdental.com, prenvalleybuilders.com, whipsaw.com, thewireguyelectric.com, quartr.com, scaleway.com, eltiempo.com, eurostar.com, france.fr, ally.com |
Which stacks put their markup where, across eighteen strata
The corpus is stratified twice over, by the platform a site is built on and by the industry it trades in, and both cuts say something the aggregate hides. The platform cut is the cleaner story about causation, because the platform is what decides whether prose and markup reach a crawler at all. 50 of the 51 Wix and Squarespace pages and 26 of the 27 WordPress pages carried a JSON-LD block in the HTML: those platforms emit it server side and the question barely arises. At the other end, 23 of the 30 static documentation sites, 28 of the 39 no code sites and 22 of the 42 Webflow sites carried none, and in most of those cases none arrived after rendering either.
The stratum where client-side assembly is most expected turned out to be the quietest. 14 of the 40 single page application startups carried no block in the HTML, and not one of the 40 gained a block from rendering, so the architecture that runs the most JavaScript was not the architecture hiding markup behind it. That is consistent with what this corpus found measuring prose rather than markup, where JavaScript supplied 7.6 percent of the words and all of it on 11 of 271 pages, and with the run where rendering added no new crawl paths. A page can be built entirely in the browser and still ship no schema in either view.
The industry cut is where the 27 concentrate. Finance and healthcare supplied 12 of them between 168 pages, against 0 from 51 news sites, and news is the industry stratum least likely to be missing markup in the HTML, with 5 of 51 carrying none. Regulated industries running large content management systems behind client-side front ends are the profile, which is the same population whose organization node is missing or unnamed in earlier runs. If you run one of these stacks, the per-stack notes for Next.js and React cover where the markup goes, and what GPTBot sees will show you your own page as a non-rendering fetch returns it.
| Stratum | Clean pairs | No JSON-LD in HTML | Gained | Added |
|---|---|---|---|---|
| industry/saas | 114 | 27 | 4 | 4 |
| industry/education | 91 | 65 | 1 | 0 |
| industry/finance | 85 | 33 | 6 | 1 |
| industry/healthcare | 83 | 45 | 6 | 0 |
| industry/government | 82 | 54 | 1 | 0 |
| industry/travel | 66 | 32 | 3 | 2 |
| industry/ecommerce | 51 | 17 | 2 | 1 |
| industry/news | 51 | 5 | 0 | 1 |
| platform/wix-squarespace | 51 | 1 | 0 | 3 |
| platform/webflow | 42 | 22 | 1 | 4 |
| platform/spa-startups | 40 | 14 | 0 | 0 |
| platform/bubble-nocode | 39 | 28 | 1 | 1 |
| platform/saas-marketing | 34 | 11 | 2 | 1 |
| platform/shopify-dtc | 31 | 9 | 0 | 0 |
| platform/static-docs | 30 | 23 | 0 | 0 |
| platform/wordpress-smb | 27 | 1 | 0 | 3 |
| platform/framer | 26 | 16 | 0 | 2 |
| platform/media-local | 26 | 1 | 0 | 2 |
What this run did not measure, and what it threw away
The largest limit is the one the figures cannot cross. No crawler was observed reading any of these pages, and no server logs were held, so nothing here supports a claim that GPTBot, ClaudeBot, PerplexityBot or any other named bot missed a block. The headless Chromium used for the second read stands in for a rendering client and is not any vendor's crawler. What this run establishes is where the markup is, which is a fact about the sites. What an engine does about it is a separate question that the methodology page is explicit about not answering, and the honest form of the finding is the conditional one: a client that does not render sees the HTML column and nothing else.
The second limit is that the exclusions were not random. 75 of the 1,044 pairs were dropped: 24 whose title matched a challenge signature, 13 of them the string "One moment, please...", 43 whose render produced a degraded document, and 8 whose raw body carried a web application firewall interstitial with fewer than 50 words of text, among them pennmedicine.org, zurich.com and gtbank.com at 212 bytes each. Those pages were served a bot wall rather than a page, and counting a bot wall as a site with no structured data would have manufactured the finding instead of measuring it. An earlier pass of this analysis did exactly that, and the correction removed 10 sites that had looked like the headline. The exclusions concentrate in strata that sit behind firewalls, so the remaining 969 under-represents heavily defended sites, which is a bias toward sites that let a crawler in.
Three narrower limits are worth stating. The render was given a fixed budget of the load event plus three seconds, so markup injected later than that was not seen, which can only make the 27 and the 25 too low rather than too high. 7 hostnames answered robots.txt with HTTP 202 and this run did not parse those bodies, so they were left out rather than measured, and that is a gap in the harness rather than a property of those sites. And 7 of the 969 produced fewer blocks after rendering than before, which would be the interesting inverse finding if it could be separated from a script failing to load behind this network; it cannot be, from one location with one attempt, so no claim is made from it. Everything above describes these 1,419 hostnames on one day, and the corpus is an editorial sampling frame assembled for coverage rather than a random draw of the web.
-
Where the markup isMeasured Two reads of the same URL differing only in whether scripts ran, on 969 hosts, with the raw read reproduced on 968 of them. -
Which crawler missed itNot measured No crawler was observed and no server logs were held. The browser used here is not any vendor's bot. -
Whether it costs citationsNot measured Nothing in this run links a block's location to an appearance in any answer engine. -
Late injected markupUndercounted The render budget was the load event plus three seconds, so anything slower was missed. This biases the counts down. -
Heavily defended sitesUnder-represented 75 pairs were excluded as challenge pages or degraded renders, and those cluster behind firewalls rather than at random. -
Markup lost to JavaScriptWithheld 7 pages produced fewer blocks after rendering, indistinguishable from a script blocked on this network, so no figure is claimed.
Lantad
Published .
A page can carry structured data in two places. It can sit in the bytes the server sends, where anything that reads HTML will find it, or it can be written into the document by JavaScript after the page loads, where only a client that executes scripts will ever see it. JavaScript structured data is the second case, and it matters here because the two populations of reader are not the same. Google says in writing that it renders pages and reads what rendering produces. Most AI crawler operators say nothing on the subject in either direction.
Common questions
Does JavaScript structured data work for AI search?
Google states that it does for Google Search: its guidance on generating structured data with JavaScript, carrying Last updated 2025-12-10 UTC, says Google Search can understand and process structured data that's available in the DOM when it renders the page. For AI crawlers the answer is undocumented rather than known. Of the 9 vendor documentation pages behind the 15 crawler tokens Lantad tracks, read on 5 September 2026, 2 stated whether their crawler executes JavaScript and 7 stated nothing either way.
Is my schema invisible because a script injects it?
Usually not, on the evidence of this run. Of 969 home pages Lantad read twice on 29 September 2026, 907 carried the identical number of JSON-LD blocks before and after a browser ran them. The pages where JavaScript was the only source of any block numbered 27, out of 404 that carried none in the HTML. The more common case by a wide margin, 374 of those 404, is a page with no structured data in either read.
Which schema types are most often added by JavaScript?
Identity and ratings, in that order. Across the 52 sites whose block count changed on 29 September 2026, Organization appeared only in the rendered DOM on 21, WebSite on 15, PostalAddress on 14 and ContactPoint on 13. The ratings cluster follows: AggregateRating on 8, Review on 5, Rating on 4, and Product and Brand on 4 each. FAQPage appeared late on 5 sites and VideoObject on 4.
How do I check whether my own structured data is in the HTML?
Request your own page without a browser and look for script elements of type application/ld+json in the bytes that come back, then load the same URL in a browser and count them again. A difference is markup that only a rendering client will ever read. Lantad's what GPTBot sees tool performs the non-rendering half of that comparison on a single URL, and its scan reports the same split across a site.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.