How readable is the web to AI crawlers?
One number over everything Lantad has measured, updated as pages are scanned. It is a live index rather than a study, and this page is explicit about the difference. The methodology is public.
Three things to take from this index.
- Every figure is measured from real scans of real pages, never modelled or estimated from a third-party index.
- The sample grows as the web is scanned, so a number quoted without its snapshot date will drift out of true.
- This is a readability measurement, not a ranking: it says what a crawler received, not what any assistant did with it.
Four figures, updated live
Every one of them is an aggregate over scans run through Lantad, with the sample size published beside it.
-
Mean AI Visibility Score
The composite across all measured pages, with the distribution across grade bands rather than the average alone, because an average hides a bimodal population.
-
Mean Prose Parity
How much of a typical page's rendered content reaches a non-rendering crawler. This is the figure that moves least and matters most.
-
Grade distribution
How many sites land in each of the five grade bands, drawn to scale below, so a mean of 75 cannot hide how it is composed.
-
The readability extremes
The share of pages that are largely invisible to AI crawlers, and the share that are fully readable. The middle is detail; the ends are the finding.
Which platforms can AI read?
This index counts every site measured, including the ones that came here and scanned themselves. That makes it the right number for how readable the measured web is, and the wrong one for how readable the web is.
So a second study samples sites chosen in advance, grouped by the platform they are built on: Framer, Webflow, Shopify, WordPress, Bubble and more. What decides whether your words reach a crawler is the thing that renders them, and unlike your industry, that is something you can change. See the crawlability study.
- Framer
- Webflow
- Wix and Squarespace
- Bubble and no-code apps
- Single-page apps
- Shopify storefronts
- WordPress
- Static site generators and docs
- SaaS marketing sites
- Media and local news
What crawlers actually get
The single biggest gap is content that never arrives: text rendered by JavaScript that most AI crawlers do not execute. Across the sample, the average page ships 70% of its rendered content in the raw HTML a crawler reads. Structured data is the weakest link.
Grade distribution
Based on 959 pages that produced a full grade. Pages that could not be graded honestly (bot walls, auth walls, non-HTML) are excluded from the shares. Snapshot: 2026-09-11.
How does your site compare?
Most pages score worse than their owners expect. Run a free scan and see exactly where yours lands, then keep it above the line as your site changes.
How the index is built
One page, two fetches
Every entry starts as a single scan: the raw HTML a crawler receives with no JavaScript, and the same page rendered in a real browser. The text shared between the two is the parity that gets graded and joins the index.
Readability is not answers
This index measures what crawlers can read. What a model actually says about your brand is a separate question, answered live against one named model on the what AI says check.
What the index measures
Four things, the same four every individual report grades, combined with the same weights:
- Prose Parity (50% of the composite). The share of the browser-rendered main content that also exists in the raw HTML a non-rendering crawler receives. This is the number that collapses when content only appears after JavaScript runs.
- Bot access (25%). Whether robots.txt rules, server-level user agent enforcement, or noai and nosnippet directives stop AI crawlers from reading the page at all.
- Structure (15%). The basics that make extracted text usable: a title, a meta description, a single H1, a canonical, regular headings, no oversized unbroken text blocks, and an llms.txt as a minor check.
- Structured data (10%). JSON-LD that parses, declares types, and carries the required properties. Structured data injected only after JavaScript runs is worth at most a quarter of this sub-score, because non-rendering crawlers never see it.
The full scoring model, including the honesty mechanisms for pages that cannot be measured cleanly, is on the methodology page.
How an entry is computed
Every entry starts life as an ordinary scan. Our fetcher requests the page once with its honest LantadBot user agent and no JavaScript, capturing exactly the HTML a non-rendering crawler receives. A real browser then renders the same URL, and the rendered main content is split into overlapping 8-token shingles and checked for containment in the crawler-side HTML. Access, structure, and structured data are scored from the same scan's evidence, and the weighted composite is graded on the same bands as every report.
A graded scan joins this index automatically: the aggregate is recomputed from the database and cached for 1 hour. It counts one row per site, using that site's most recent graded scan, and it leaves out pages on our own domain along with the sources that would count one site many times, which are weekly monitor re-scans and the individual pages of a multi-page or batch run. Nothing is estimated from third-party indexes and nothing is hand-curated. A page you scan today is part of the sample as soon as it is graded and the hourly cache refreshes.
What this sample is, and is not
- Not a random sample. Everyone here chose to scan. That skews toward people who suspected a problem, which almost certainly makes the mean lower than the web's true mean. A growing share comes from a corpus we scan ourselves, curated by hand, so that is not a random crawl either.
- One row per site. Each site counts once, through the most recent graded scan of it, however many times it was scanned. An entry is a single URL fetched as a visitor and as a crawler.
- Not a ranking. No site is named without consent. The index publishes distributions, not a leaderboard of who is failing.
- Graded scans only. Pages behind bot walls or auth walls and non-HTML responses produce no composite score and are excluded, so the sample understates how much of the web cannot be measured at all.
- Readability, not answers. The index measures what crawlers can read, not what any assistant currently says about a brand. For that question on your own page, run the what AI says check, which asks one named open-weight model live.
Numbers appear on this page only once at least 5 graded scans exist. An average over a handful of pages is an anecdote, not a survey.
How to cite this
The figures move as more pages are measured, so a citation needs a date.
-
Quote the snapshot date
Every figure on the live page carries the date it was read. Cite the number and the date together, because the number will change.
-
Link the page
Licensed CC BY 4.0, which asks for attribution and a link. Nothing else is required and nothing is behind a form.
-
Say which dataset
The live index and the crawlability study measure different populations and answer different questions. Name the one you used.
This index counts everyone who scanned, and says so. Sites that ran a scan on Lantad chose to be measured, which makes them a different population from the web at large. The index reports what our visitors' pages score. The crawlability study reports a seeded sample chosen in advance, and excludes self-selected scans entirely.
Both are published openly under CC BY 4.0. Neither publishes a figure before it has enough measurements to support one.
Common questions
How often does the index update?
Continuously. The aggregate is recomputed from the live database and cached for 1 hour, so a newly scanned site is part of these numbers within the hour. The snapshot date under the grade chart is the day the page was rendered; the aggregate behind it is at most an hour older.
What counts as one site in the sample?
One site, counted once, using its most recent scan that produced a full composite score. Earlier scans of the same site are not counted again, so a site we have scanned fifty times still counts once and its current state is what appears. www and the bare domain are the same site. Our own domain is excluded, and so is anything that could not be graded honestly: bot walls, auth walls, non-HTML responses. Repeat scans driven by monitoring, and the individual pages of a whole-site scan, are excluded from the sample as well: counting them would describe our customers' sites rather than the web. Sites we scanned ourselves as part of the curated corpus are included and counted the same way as any other: one site, one scan, its homepage.
Which AI crawlers does readability refer to?
The crawler view is a plain HTTP fetch with no JavaScript execution, which is how most AI crawlers request pages. Each scan also evaluates robots.txt for 15 published AI crawler tokens, from GPTBot to Amazonbot, and probes each one that publishes a request user agent of its own.
Can I cite these numbers?
Yes, with a link to this page. The methodology is public, and the sample grows as the web is scanned, so the numbers move; quote the snapshot date shown with the chart alongside any figure.
Why is the mean score lower than I expected?
Because of who is in it. People scan when they suspect a problem, so this population is skewed toward sites with one. The crawlability study exists precisely because a self-selected sample cannot answer what the web scores.
Can I get the raw data?
Aggregates are published on the page under CC BY 4.0. Per-site raw data is not released, because a scan report belongs to the person who ran it.
Do you name sites?
Not without consent. Published research names a site only where the owner agreed or where the finding is already public and the site is being credited rather than criticised.
Free to reuse and republish with attribution and a link to this page. The index is published under CC BY 4.0.
Add your own page to the picture.
Every scan feeds the index, and the index is the only reason anyone can say what a normal parity score looks like.