Home / Research / Crawlability study

Research

Which platforms can AI crawlers read?

A stratified sample of sites, grouped by the platform they are built on, scanned one page each. The question it answers is the one the live index cannot: whether your kind of site is the broken kind. Every figure carries the number of sites behind it.

10 platform strata
392 sites seeded
378 sites measured

96% of the frame measured

Snapshot 2026-09-22 · scanned a few sites a night

Headline figures

What the sample shows

81 mean AI Visibility Score across 378 sampled sites
79% of rendered content reaches a crawler on the average sampled page
11% ship almost nothing to crawlers (under 30% prose parity)

Grade distribution

Every measured site in the sample, by the grade its score falls into.

119 A
165 B
47 C
31 D
16 F

By platform

Sorted by score, worst first, so the platforms with the most to fix lead. A row publishes figures only once at least 5 of its sites have been measured.

Platform Measured Mean score Mean parity Near invisible
Bubble and no-code apps 40 / 42 72 67% 28%
Webflow 41 / 42 75 73% 12%
Shopify storefronts 31 / 34 79 73% 6%
Single-page apps 45 / 45 79 74% 18%
SaaS marketing sites 38 / 39 80 75% 13%
Media and local news 30 / 32 84 81% 7%
WordPress 40 / 41 84 82% 13%
Framer 32 / 32 85 90% 0%
Static site generators and docs 31 / 31 85 95% 3%
Wix and Squarespace 50 / 54 86 87% 8%
Sampling

How this was sampled

The frame was chosen before anything was scanned, which is the part that makes it a study rather than a summary of our traffic. Sites were drawn from public showcases, directories and platform galleries, 10 strata in all, then scanned a few per night in the order they were listed. The sources each stratum came from are listed below.

  1. 54 public lists

    Showcases, directories and platform galleries, every one of them named below so the frame can be checked rather than trusted.

  2. 10 platform strata

    Sorted by the thing that renders the page, because that is what decides whether your words reach a crawler.

  3. 392 sites seeded

    One page per site, its home page, counted once. The list was fixed before anything was fetched.

  4. 378 sites measured

    Scanned a few per night in the order they were listed, so the table above fills in over weeks rather than all at once.

Why platform, not industry

Splitting by platform rather than by industry is the whole point. What decides whether your words reach a crawler is the thing that renders them, not what your business sells: a Framer marketing site and a Bubble app fail the same way as each other and a completely different way from a WordPress blog. It is also the only split a reader can act on, because you can change how a page is built and you cannot change what industry you are in.

What a stratum claims

The stratum is still a judgement. Identifying a builder from the outside is inference, and a site can be rebuilt after it was listed, so a row describes the sites we filed under it rather than every site on that platform.

One site, one page, counted once

One page per site, its home page, counted once. A site that is re-measured later replaces its earlier result rather than adding a second row. Sites that came to Lantad and scanned themselves are excluded from every figure on this page; they are counted in the live index, which says so.

Only where robots.txt allows

Sites are only fetched where robots.txt allows our crawler, and a site that declines is skipped and not asked again for a month. The conduct rules are on /bot, and the scoring method is on /methodology.

A limitation worth stating

Our renderer reads the page's document. Text held inside shadow DOM, which some component frameworks use, is not in that document, so a site built that way would measure lower than it deserves. That would matter here more than anywhere, because a whole platform could be marked down for a rendering choice rather than for what a crawler actually receives.

So we measured it rather than assuming. Across a random sample of the scanned corpus, no page used declarative shadow DOM and none called attachShadow: the only custom elements found were behaviour wrappers from a Shopify theme, with their text in the ordinary document. On this frame the effect is nil. It is stated because the sample is not the whole web, and a platform that leans on shadow DOM would need this fixed before it could be fairly compared.

Provenance

Where each stratum came from

Counts are sites seeded, not sites measured. Published so the frame can be audited rather than taken on trust.

  • Bubble and no-code apps

    42sites seeded

    • shno.co/blog/bubble-app-examples
    • shno.co/blog/carrd-websites
    • shno.co/blog/glide-app-examples
    • shno.co/blog/softr-app-examples
    • softr.io/customer-stories/altitude-marketing
    • softr.io/customer-stories/eight-digit-media
    • softr.io/customer-stories/lakeshore-windows
    • softr.io/customer-stories/no-code-week
    • softr.io/customer-stories/the-board
  • Webflow

    42sites seeded

    • createtoday.io/examples?category=community&platform=webflow
    • createtoday.io/examples?category=education&platform=webflow
    • createtoday.io/examples?category=media&platform=webflow
    • createtoday.io/examples?platform=webflow
    • flowout.com/portfolio
    • flowzai.com/blog-post/best-webflow-websites
    • htmlburger.com/blog/webflow-ecommerce-examples/
    • joinamply.com/post/best-webflow-websites
    • todaymade.com/blog/websites-built-with-webflow
    • wedoflow.com/post/webflow-success-stories-popular-and-impactful-websites-built-with-webflow
  • Shopify storefronts

    34sites seeded

    • ecommerceparadise.com/50-best-shopify-stores-in-2026
    • fastbundle.co/blog/best-shopify-stores
    • shopify.com/blog/niche-stores
    • sitebuilderreport.com/inspiration/shopify-niche-stores
  • Single-page apps

    45sites seeded

    • extruct.ai/ycombinator-companies/w25
    • producthunt.com/leaderboard/monthly/2025/11
    • producthunt.com/leaderboard/monthly/2025/6
    • producthunt.com/leaderboard/monthly/2025/8
    • producthunt.com/leaderboard/monthly/2026/1
    • producthunt.com/leaderboard/monthly/2026/3
    • producthunt.com/leaderboard/monthly/2026/4
    • producthunt.com/leaderboard/monthly/2026/5
  • SaaS marketing sites

    39sites seeded

    • failory.com/startups/developer-tools
    • failory.com/startups/fintech
    • failory.com/startups/health-care
    • failory.com/startups/human-resources
    • failory.com/startups/logistics
    • failory.com/startups/marketing
  • Media and local news

    32sites seeded

    • lionpublishers.com/lion-welcomes-new-members-from-18-states/
    • lionpublishers.com/meet-the-51-finalists-for-the-2025-lion-sustainability-awards/
    • web search
    • www.anoffgridlife.com/homestead-blogs/
  • WordPress

    41sites seeded

    • delmain.co/blog/best-dental-websites/
    • web search
    • www.allianceinteractive.com/50-best-accounting-website-examples/
  • Framer

    32sites seeded

    • brixtemplates.com/blog/best-framer-agencies
    • goodspeed.studio/framer-website-examples/e-commerce
    • landing.gallery/website-builder/framer
    • sitebuilderreport.com/inspiration/framer-websites
  • Static site generators and docs

    31sites seeded

    • docusaurus.io/showcase
    • www.11ty.dev/speedlify/
  • Wix and Squarespace

    54sites seeded

    • createtoday.io/examples?category=beauty-salon&platform=squarespace
    • sitebuilderreport.com/inspiration/local-business-websites
    • sitebuilderreport.com/inspiration/squarespace-business-websites
    • sitebuilderreport.com/inspiration/squarespace-charity-websites
    • sitebuilderreport.com/wix-examples
    • tooltester.com/en/blog/wix-website-examples/
    • web search
The method

What is measured, and how

One page per site, its home page, fetched twice and diffed. The full method is on the methodology page.

  • Two fetches, one diff

    Once as an AI crawler sees it, with no JavaScript, and once in a real browser. Prose Parity is how much of the rendered content survives the first fetch.

  • Access, per token

    Each scan also evaluates robots.txt for all 15 published AI crawler tokens, so a stratum's access profile is measured rather than assumed.

  • A minimum sample per row

    No row publishes a mean before enough of its seeded sites have been measured. Rows below the floor show their real count and nothing else.

  • Platform by sampling frame

    Each stratum was drawn from that platform's own showcases and galleries. It is an inference from how a site was sampled, not a fingerprint of the site.

The sample was chosen, not self-selected. Sites were seeded in advance from public showcases, directories and platform galleries across 10 strata, and are scanned a few per night. Sites that scanned themselves on Lantad are deliberately excluded, because they chose to be measured and mixing the two would describe our visitors rather than the web.

The drip runs slowly, so for the first weeks most rows read not enough yet with their real count. That is the page working, not the page broken.

Common questions

Which sites are in this study?

A seeded list of 392 sites across 10 platform strata, chosen in advance from public showcases, directories and platform galleries, and scanned a few per night. Sites that scanned themselves on Lantad are deliberately NOT counted here: they chose to be measured, which makes them a different population, and mixing the two would describe our visitors rather than the web. The wider live index at /research does count them and says so.

Why do some rows say there is not enough data?

Because there is not. A row publishes a figure only once at least 5 of its seeded sites have been measured, and the background scan adds a few sites a night, so strata fill up over weeks. The alternative is a mean over two sites presented as a fact about a whole platform, which is the sort of number this product exists to argue against.

Why split by platform rather than by industry?

Because the platform is what decides the answer. Whether your words reach an AI crawler is determined by the thing that renders them, so a Framer marketing site and a Bubble app fail the same way as each other and a completely different way from a WordPress blog, however unrelated their businesses are. Industry would group sites by something that has nothing to do with why they score what they score. Platform is also the only split a reader can act on: you can change how a page is built, and you cannot change what industry you are in.

How is a site's platform decided?

From how it was sampled: each stratum was drawn from public showcases, directories and galleries for that platform, and those sources are listed on this page. It remains an inference rather than a measurement, and a site can be rebuilt after it was listed, so a row describes the sites we filed under it rather than every site on that platform.

What exactly is measured?

One page per site, its home page, fetched twice: once as an AI crawler sees it, with no JavaScript, and once in a real browser, then diffed. Prose parity is how much of the rendered content survives that first fetch. Each scan also evaluates robots.txt for 15 published AI crawler tokens. The full method is on /methodology.

Can I cite this?

Yes, with a link to this page. The figures move as more of the frame is measured, so quote the snapshot date shown beside them. Licensed CC BY 4.0.

Free to reuse and republish with attribution and a link to this page. These figures are published under CC BY 4.0. The sampling frame behind them is described above.

Your platform is a tendency, not a verdict.

A stratum tells you what sites like yours usually score. Only a scan tells you what yours does.