Home / Research / Crawlability study
ResearchWhich platforms can AI crawlers read?
A stratified sample of sites, grouped by the platform they are built on, scanned one page each. The question it answers is the one the live index cannot: whether your kind of site is the broken kind. Every figure carries the number of sites behind it.
96% of the frame measured
What the sample shows
Grade distribution
Every measured site in the sample, by the grade its score falls into.
By platform
Sorted by score, worst first, so the platforms with the most to fix lead. A row publishes figures only once at least 5 of its sites have been measured.
| Platform | Measured | Mean score | Mean parity | Near invisible |
|---|---|---|---|---|
| Bubble and no-code apps | 40 / 42 | 72 | 67% | 28% |
| Webflow | 41 / 42 | 75 | 73% | 12% |
| Shopify storefronts | 31 / 34 | 79 | 73% | 6% |
| Single-page apps | 45 / 45 | 79 | 74% | 18% |
| SaaS marketing sites | 38 / 39 | 80 | 75% | 13% |
| Media and local news | 30 / 32 | 84 | 81% | 7% |
| WordPress | 40 / 41 | 84 | 82% | 13% |
| Framer | 32 / 32 | 85 | 90% | 0% |
| Static site generators and docs | 31 / 31 | 85 | 95% | 3% |
| Wix and Squarespace | 50 / 54 | 86 | 87% | 8% |
How this was sampled
The frame was chosen before anything was scanned, which is the part that makes it a study rather than a summary of our traffic. Sites were drawn from public showcases, directories and platform galleries, 10 strata in all, then scanned a few per night in the order they were listed. The sources each stratum came from are listed below.
-
54
public lists
Showcases, directories and platform galleries, every one of them named below so the frame can be checked rather than trusted.
-
10
platform strata
Sorted by the thing that renders the page, because that is what decides whether your words reach a crawler.
-
392
sites seeded
One page per site, its home page, counted once. The list was fixed before anything was fetched.
-
378
sites measured
Scanned a few per night in the order they were listed, so the table above fills in over weeks rather than all at once.
Why platform, not industry
Splitting by platform rather than by industry is the whole point. What decides whether your words reach a crawler is the thing that renders them, not what your business sells: a Framer marketing site and a Bubble app fail the same way as each other and a completely different way from a WordPress blog. It is also the only split a reader can act on, because you can change how a page is built and you cannot change what industry you are in.
What a stratum claims
The stratum is still a judgement. Identifying a builder from the outside is inference, and a site can be rebuilt after it was listed, so a row describes the sites we filed under it rather than every site on that platform.
One site, one page, counted once
One page per site, its home page, counted once. A site that is re-measured later replaces its earlier result rather than adding a second row. Sites that came to Lantad and scanned themselves are excluded from every figure on this page; they are counted in the live index, which says so.
Only where robots.txt allows
Sites are only fetched where robots.txt allows our crawler, and a site that declines is skipped and not asked again for a month. The conduct rules are on /bot, and the scoring method is on /methodology.
A limitation worth stating
Our renderer reads the page's document. Text held inside shadow DOM, which some component frameworks use, is not in that document, so a site built that way would measure lower than it deserves. That would matter here more than anywhere, because a whole platform could be marked down for a rendering choice rather than for what a crawler actually receives.
So we measured it rather than assuming. Across a random sample of the scanned corpus, no page used declarative shadow DOM and none called attachShadow: the only custom elements found were behaviour wrappers from a Shopify theme, with their text in the ordinary document. On this frame the effect is nil. It is stated because the sample is not the whole web, and a platform that leans on shadow DOM would need this fixed before it could be fairly compared.
Where each stratum came from
Counts are sites seeded, not sites measured. Published so the frame can be audited rather than taken on trust.
-
Bubble and no-code apps
42sites seeded
-
shno.co/blog/bubble-app-examples -
shno.co/blog/carrd-websites -
shno.co/blog/glide-app-examples -
shno.co/blog/softr-app-examples -
softr.io/customer-stories/altitude-marketing -
softr.io/customer-stories/eight-digit-media -
softr.io/customer-stories/lakeshore-windows -
softr.io/customer-stories/no-code-week -
softr.io/customer-stories/the-board
-
-
Webflow
42sites seeded
-
createtoday.io/examples?category=community&platform=webflow -
createtoday.io/examples?category=education&platform=webflow -
createtoday.io/examples?category=media&platform=webflow -
createtoday.io/examples?platform=webflow -
flowout.com/portfolio -
flowzai.com/blog-post/best-webflow-websites -
htmlburger.com/blog/webflow-ecommerce-examples/ -
joinamply.com/post/best-webflow-websites -
todaymade.com/blog/websites-built-with-webflow -
wedoflow.com/post/webflow-success-stories-popular-and-impactful-websites-built-with-webflow
-
-
Shopify storefronts
34sites seeded
-
ecommerceparadise.com/50-best-shopify-stores-in-2026 -
fastbundle.co/blog/best-shopify-stores -
shopify.com/blog/niche-stores -
sitebuilderreport.com/inspiration/shopify-niche-stores
-
-
Single-page apps
45sites seeded
-
extruct.ai/ycombinator-companies/w25 -
producthunt.com/leaderboard/monthly/2025/11 -
producthunt.com/leaderboard/monthly/2025/6 -
producthunt.com/leaderboard/monthly/2025/8 -
producthunt.com/leaderboard/monthly/2026/1 -
producthunt.com/leaderboard/monthly/2026/3 -
producthunt.com/leaderboard/monthly/2026/4 -
producthunt.com/leaderboard/monthly/2026/5
-
-
SaaS marketing sites
39sites seeded
-
failory.com/startups/developer-tools -
failory.com/startups/fintech -
failory.com/startups/health-care -
failory.com/startups/human-resources -
failory.com/startups/logistics -
failory.com/startups/marketing
-
-
Media and local news
32sites seeded
-
lionpublishers.com/lion-welcomes-new-members-from-18-states/ -
lionpublishers.com/meet-the-51-finalists-for-the-2025-lion-sustainability-awards/ - web search
-
www.anoffgridlife.com/homestead-blogs/
-
-
WordPress
41sites seeded
-
delmain.co/blog/best-dental-websites/ - web search
-
www.allianceinteractive.com/50-best-accounting-website-examples/
-
-
Framer
32sites seeded
-
brixtemplates.com/blog/best-framer-agencies -
goodspeed.studio/framer-website-examples/e-commerce -
landing.gallery/website-builder/framer -
sitebuilderreport.com/inspiration/framer-websites
-
-
Static site generators and docs
31sites seeded
-
docusaurus.io/showcase -
www.11ty.dev/speedlify/
-
-
Wix and Squarespace
54sites seeded
-
createtoday.io/examples?category=beauty-salon&platform=squarespace -
sitebuilderreport.com/inspiration/local-business-websites -
sitebuilderreport.com/inspiration/squarespace-business-websites -
sitebuilderreport.com/inspiration/squarespace-charity-websites -
sitebuilderreport.com/wix-examples -
tooltester.com/en/blog/wix-website-examples/ - web search
-
What is measured, and how
One page per site, its home page, fetched twice and diffed. The full method is on the methodology page.
-
Two fetches, one diff
Once as an AI crawler sees it, with no JavaScript, and once in a real browser. Prose Parity is how much of the rendered content survives the first fetch.
-
Access, per token
Each scan also evaluates robots.txt for all 15 published AI crawler tokens, so a stratum's access profile is measured rather than assumed.
-
A minimum sample per row
No row publishes a mean before enough of its seeded sites have been measured. Rows below the floor show their real count and nothing else.
-
Platform by sampling frame
Each stratum was drawn from that platform's own showcases and galleries. It is an inference from how a site was sampled, not a fingerprint of the site.
The sample was chosen, not self-selected. Sites were seeded in advance from public showcases, directories and platform galleries across 10 strata, and are scanned a few per night. Sites that scanned themselves on Lantad are deliberately excluded, because they chose to be measured and mixing the two would describe our visitors rather than the web.
The drip runs slowly, so for the first weeks most rows read not enough yet with their real count. That is the page working, not the page broken.
Common questions
Which sites are in this study?
A seeded list of 392 sites across 10 platform strata, chosen in advance from public showcases, directories and platform galleries, and scanned a few per night. Sites that scanned themselves on Lantad are deliberately NOT counted here: they chose to be measured, which makes them a different population, and mixing the two would describe our visitors rather than the web. The wider live index at /research does count them and says so.
Why do some rows say there is not enough data?
Because there is not. A row publishes a figure only once at least 5 of its seeded sites have been measured, and the background scan adds a few sites a night, so strata fill up over weeks. The alternative is a mean over two sites presented as a fact about a whole platform, which is the sort of number this product exists to argue against.
Why split by platform rather than by industry?
Because the platform is what decides the answer. Whether your words reach an AI crawler is determined by the thing that renders them, so a Framer marketing site and a Bubble app fail the same way as each other and a completely different way from a WordPress blog, however unrelated their businesses are. Industry would group sites by something that has nothing to do with why they score what they score. Platform is also the only split a reader can act on: you can change how a page is built, and you cannot change what industry you are in.
How is a site's platform decided?
From how it was sampled: each stratum was drawn from public showcases, directories and galleries for that platform, and those sources are listed on this page. It remains an inference rather than a measurement, and a site can be rebuilt after it was listed, so a row describes the sites we filed under it rather than every site on that platform.
What exactly is measured?
One page per site, its home page, fetched twice: once as an AI crawler sees it, with no JavaScript, and once in a real browser, then diffed. Prose parity is how much of the rendered content survives that first fetch. Each scan also evaluates robots.txt for 15 published AI crawler tokens. The full method is on /methodology.
Can I cite this?
Yes, with a link to this page. The figures move as more of the frame is measured, so quote the snapshot date shown beside them. Licensed CC BY 4.0.
Free to reuse and republish with attribution and a link to this page. These figures are published under CC BY 4.0. The sampling frame behind them is described above.
Your platform is a tendency, not a verdict.
A stratum tells you what sites like yours usually score. Only a scan tells you what yours does.