BlogFindings
Internal links and AI crawlers: 84 of 1,091 home pages offered a crawler no path at all
Lantad requested the home page of all 1,419 hostnames in this repository's committed corpus on 21 September 2026 and read the raw bytes with no JavaScript executed. 1,091 answered HTTP 200 with HTML. 84 of those carried no internal link a crawler could follow, and 55 of the 84 contained no anchor element at all. On the 1,007 that did offer a path, the first one was requested too.
So this run asked a narrower question than usual. Not what the home page says, but what it offers: how many distinct paths into the rest of the site are present in the raw HTML, and what is at the end of the first one. The answer on 21 September 2026 is that most sites in this corpus are generous with paths, the median offering 48 of them, and that a minority offer a crawler nothing at all. The names in that minority are not the ones you would guess.
In short
- Internal links are the only way an AI crawler reaches a page other than the one it arrived on, and 84 of the 1,091 home pages that answered Lantad on 21 September 2026 carried none that a crawler could follow.
- 55 of those 84 home pages contained no anchor element at all in the raw HTML, among them ryanair.com, nordstrom.com, schwab.com and archive.org, each re-fetched and re-confirmed on the same day.
- Google's Make your links crawlable documentation, carrying Last updated 2025-12-10 UTC, states that Google can only crawl a link if it is an a HTML element with an href attribute, which is the rule this measurement counted against.
- Of the 604 home pages carrying JSON-LD on 21 September 2026, 143 led to a first internal page that carried none, so structured data did not survive one click on 23.7 percent of the sites that had it.
- The first internal link returned HTTP 404 on 18 of the 1,007 home pages that offered one, and on eight of those 18 the href held an unrendered template expression or a fragment of JavaScript rather than a URL.
| Stage | Sites | What happened |
|---|---|---|
| Hostnames asked | 1,419 | The committed corpus, an editorial frame rather than a random draw |
| Never returned a status | 31 | 19 failed DNS resolution, 11 timed out or aborted, 1 failed in the client |
| Returned a status | 1,388 | A response arrived, of any kind |
| Refused this crawler | 214 | Answered HTTP 403 to LantadBot |
| Answered HTTP 503 | 55 | Service unavailable, so nothing to read |
| Answered 200 with HTML | 1,091 | The denominator for every figure below |
| Offered no internal link | 84 | No path into the rest of the site in the raw bytes |
| Offered at least one | 1,007 | The first one in document order was then requested |
What counts as an internal link an AI crawler can follow
A link is a thing a crawler can act on, not a thing a user can click. Those two sets overlap and they are not the same set, and the difference is where this measurement lives.
Google's Make your links crawlable documentation, carrying Last updated 2025-12-10 UTC when it was opened on 21 September 2026, is unusually direct about the rule. It says that generally Google can only crawl a link if it is an a HTML element with an href attribute, that most links in other formats will not be parsed and extracted by its crawlers, and that it cannot reliably extract URLs from a elements without an href or from other tags that perform as links because of script events. The same page says Google uses links for page relevancy and discovery, which is the reason any of this matters.
The rule follows from what the element is. The HTML anchor element, as MDN defines it, creates a hyperlink with its href attribute, to web pages, files, email addresses, locations in the same page, or anything else a URL can address. Three of those destinations are not another page on this host, which is why the count below excludes them one by one rather than counting anchors.
That is the rule this scan counted against. For each home page the raw HTML was searched for a elements carrying an href, each value was resolved against the final URL after redirects, and what remained after four exclusions was counted: off-host targets, fragments that point back into the same document, non-http schemes such as mailto and tel, and file extensions that are not pages. A link back to the home page itself was excluded too, since it offers no new path. What is left is the set of distinct further documents on the same host that a crawler reading these bytes could ask for next. The vocabulary underneath all of it is set out under AI crawler.
One limit belongs here rather than at the end, because it shapes every number that follows. No JavaScript was executed, and Google's page says plainly that links inserted dynamically by JavaScript are crawlable as long as they use the same markup. So a count of zero here means zero paths for a client that does not render, which is the condition several of the crawlers in the crawler directory publish for themselves and which this blog measured directly when JavaScript added no new crawl paths on a smaller set of captured pages. It is a real reading of the web for some crawlers and an incomplete one for others.
-
a element with hrefCrawlable The only form the page says Google can generally crawl. -
a element inserted by JavaScriptCrawlable Stated as crawlable provided it uses the same markup. Not visible to a client that does not render. -
a element without hrefNot reliable Named as something Google cannot reliably extract a URL from. -
Other tags acting as linksNot reliable Named alongside the above, where the behaviour comes from script events. -
Most other link formatsNot parsed The page says most links in other formats will not be parsed and extracted.
How many internal links does a home page give an AI crawler?
Across the 1,091 home pages that answered HTTP 200 with HTML, the median count is 48 distinct internal paths and the mean is 75. The gap between those two numbers is the shape of the distribution: a long tail of very large navigations pulls the average up, and the busiest single page in the corpus offers 1,573 distinct internal URLs.
Most of the corpus is comfortable. 535 of the 1,091 home pages offer 50 or more paths, and 255 offer 100 or more. For those sites the question this post asks is not interesting, because a crawler arriving at the home page has more to do than it will get through in one visit, and the useful question moves on to which of those paths are worth following. That is the territory of the sitemap declarations 839 of 1,016 robots.txt files carry, and of what happens when a crawler walks a site properly, which this blog measured when it followed 985 links on one site against 200,712 URLs declared in its sitemaps.
The thin end is where the finding is. 193 of the 1,091 home pages offer nine paths or fewer, and 84 of those offer none. A site with four internal links in its home page HTML is not necessarily a small site; it is a site where a crawler that does not render has four doors, and everything behind the other doors is reachable only from a sitemap, an external link, or a rendering pass the operator has no way to require. This is the same structural problem as prose parity and it is measured the same way: by asking what arrived, not what was authored.
One thing the count deliberately does not measure is quality. A home page offering 300 paths of which 290 are faceted product filters is worse served than one offering 30 substantive pages, and this scan cannot tell those apart. Nor does it weigh anchor text, though the corpus has shown before that anchor text goes missing at scale: 85 links across five captured pages carried no anchor text at all, 34 of them named only by an aria-label.
The 84 home pages that offered no path at all
A count of zero is the kind of result that is usually a bug in the counter, so all 84 were requested again the same day and classified by what their anchors actually did. None of them gained an internal path on the second request. The 84 split into five shapes, and they are not equally interesting.
55 of the 84 contained no anchor element whatsoever in the raw HTML. Not an anchor pointing somewhere unhelpful: none at all. The list includes ryanair.com, southwest.com, aircanada.com, britishairways.com and ethiopianairlines.com; nordstrom.com, lowes.com, jd.com, shopee.sg and takealot.com; schwab.com, rbc.com, zurich.com and bradesco.com.br; and archive.org, jstor.org, orcid.org, openstax.org and khanacademy.org. These are large organisations with large sites, and on 21 September 2026 a crawler that does not execute JavaScript could see exactly one of their pages. Most of them also returned no readable words, which is the same failure this blog counted when 17 of 380 home pages sent a crawler zero words and measured directly when JavaScript supplied all the prose on 11 of 271 pages.
16 had anchors that all pointed off the host. A personal site whose only links go to a newsletter platform, a payment host and two social profiles is complete from a reader's point of view and a dead end from a crawler's. 5 were one-page sites whose navigation is entirely fragment links into the same document, which is a legitimate design and still leaves a crawler with one URL. 4 pointed only back at the home page. The last 4 are the case Google's documentation names explicitly: anchors carrying no href a crawler can resolve to a page, including one site with 11 anchor elements and not a single href among them, and one whose only three anchors are mailto and tel links.
The distinction matters for what a site owner should do about it. A one-page site has nothing to fix and nothing to gain. A site whose navigation is built from elements that respond to script events is losing paths it believes it has, and the fix is markup rather than content, which is the shape of the per-stack advice in the React guide.
| What the anchors did | Sites | Named examples from the corpus |
|---|---|---|
| No anchor element at all | 55 | ryanair.com, nordstrom.com, schwab.com, archive.org, khanacademy.org |
| All anchors point off this host | 16 | hse.ie, visa.com, newsletteros.com, dailyui.co |
| All anchors are fragments in the same page | 5 | herdmoments.com, americanelitepainters.com, calltree.ai |
| All anchors point back at the home page | 4 | byword.ai, letterhunt.co, listr.pro, contpaqi.com |
| Anchors carry no href a crawler can resolve to a page | 4 | pollyreach.ai, agoda.com, citi.com, patagonia.com |
Which kinds of site left a crawler with nowhere to go
The corpus is stratified, so the 84 can be attributed rather than merely listed, and the pattern is clean enough to be worth stating plainly: this is a property of how the page is built, not of how large the organisation is.
The no-code stratum is the extreme. 18 of the 42 Bubble-built sites that answered, 42.9 percent, offered a crawler no internal path. The single page application startups follow at 6 of 45. Among the industry strata the rates are lower and the absolute numbers larger, because those strata are larger: travel at 12 of 78, ecommerce at 10 of 66, finance at 10 of 94 and healthcare at 9 of 95.
The other end of the table is the part worth dwelling on, because it is the good news and it is unanimous. Not one of the 54 Wix and Squarespace sites, the 37 WordPress sites, the 34 Shopify storefronts or the 38 SaaS marketing sites offered zero paths. Four platform strata, 163 sites, no failures. Those are the platforms that render navigation on the server as a matter of course, and their users did nothing to earn the result. This is the mirror image of a pattern this blog keeps finding, where a platform default decides the outcome and the site owner never sees the decision: it was true of the 141 of 382 home pages carrying no structured data and it is true here, with the sign reversed.
Two cautions on reading the table. The strata are sampling frames of unequal size and construction, so a rate is a statement about these sites rather than about the platform in general, and the smallest strata carry the least weight. And 214 hostnames refused this crawler outright with HTTP 403 before any of this could be measured, weighted heavily towards news, which is the same refusal pattern reported when 79 of 115 sites refused a crawler their robots.txt allows. Those sites are absent from every figure here.
| Stratum | No path | Answered 200 | Rate |
|---|---|---|---|
| Bubble, no-code | 18 | 42 | 42.9% |
| Travel | 12 | 78 | 15.4% |
| Ecommerce | 10 | 66 | 15.2% |
| Single page app startups | 6 | 45 | 13.3% |
| Finance | 10 | 94 | 10.6% |
| Healthcare | 9 | 95 | 9.5% |
| Government | 5 | 88 | 5.7% |
| Education | 6 | 106 | 5.7% |
| News | 1 | 61 | 1.6% |
| SaaS | 1 | 120 | 0.8% |
| Wix and Squarespace | 0 | 54 | 0.0% |
| Shopify storefronts | 0 | 34 | 0.0% |
| WordPress | 0 | 37 | 0.0% |
| SaaS marketing sites | 0 | 38 | 0.0% |
What happened when the first internal link was requested
For each of the 1,007 home pages that offered a path, the first one in document order was requested the same way. First in document order is not most important, and nothing here claims otherwise: it is simply the link a crawler reading top to bottom meets first, which makes it a fair sample of the quality of what is on offer rather than a judgement about any site's navigation.
1,006 returned a status and 969 of those answered HTTP 200. The 37 that did not divide into refusals and mistakes. 11 answered 403, which is a site treating the second request from a crawler differently from the first. 2 answered 429, 2 answered 503, and one each answered 400, 405, 406 and 500. The remaining 18 answered 404, and that is the number worth explaining, because a broken link is ordinary and most of these are not.
On six of the 18, the href was never a URL. hoshinoresorts.com carries a Vue template expression, locale.texts indexed by hotel id and wrapped in double braces, sitting in an href attribute that was never evaluated. botanicapaper.com and oaklandish.com carry a fragment of JavaScript beginning window.location.href.replace. acc.org offers CTA.url and slate.com offers desktopMastheadUrl, which are variable names that reached the client as literal paths. thetinypod.com has an escaped quote and an image host inside the attribute. Each is a template that did not render, and a crawler has no way to know that: it reads an href, resolves it, and asks for it.
A further five of the 18 are Cloudflare's email protection path, which is a redirect endpoint rather than a page and answers 404 to a client arriving cold. Two more resolve to paths that are not there: wien.gv.at offers a link to nojs, and the first link on destatis.de repeats its own path prefix, which is the signature of a relative URL resolved against the wrong base.
That leaves five ordinary broken links, including one storefront whose first internal link is a collection page that no longer exists. Ordinary is the right word, and five is the honest size of it. The point of separating them is that a 404 rate of 1.8 percent on first links sounds like link rot and mostly is not: 13 of the 18 are markup describing a link the page never actually had, which is the same class of problem as a URL that answers 200 with nothing behind it, counted yesterday in the 125 sites that served a soft 404.
first internal link, resolved and requested
- hoshinoresorts.com href="{{locale.texts['reservation.hd.faq.url.' + hotel.hotelId]}}" 404
- botanicapaper.com href="window.location.href.replace(/..." 404
- oaklandish.com href="window.location.href.replace(/..." 404
- acc.org href="CTA.url" 404
- slate.com href="desktopMastheadUrl" 404
- thetinypod.com href=""//imgur.com/a/ieO3opW"" 404
- wien.gv.at href="/nojs" 404, a real path that is not there
- destatis.de href="/DE/Home/DE/Home/_inhalt.html" 404, prefix repeated
- five further sites href="/cdn-cgi/l/email-protection" 404, a redirect endpoint
What the second page lost: prose and structured data
969 pairs of home page and first internal page can be compared directly, and on the whole the second page holds up. The median home page carried 1,009 readable words and the median first internal page carried 775, which is the shape you would expect when a home page is a summary of everything and an inner page is one thing.
Two minorities matter more than the median. 126 of the 969 inner pages carried under a quarter of their home page's word count where the home page had at least 100 words, and 13 carried no readable words at all. The 13 include starbucks.com, coinbase.com, flipkart.com, progressive.com and carnival.com, all of which returned a readable home page and then an empty second document. A site can pass a home page check and fail one click later, and nothing in a home page audit would show it. That is an argument for checking more than one URL, which is what multiscan exists to do.
Structured data travels worse than prose. 604 of the 969 home pages carried at least one JSON-LD block, and on 143 of those 604 the first internal page carried none: 23.7 percent of the sites that had it lost it one click in. The reverse happened too, on a smaller scale. 38 of the 365 home pages carrying no JSON-LD led to an inner page that had some, usually a product or article page where the platform emits markup the home page template does not.
This is worth taking seriously because of what structured data is for. It names the entity, and an answer engine that cannot resolve what a page is about has to infer it from prose, which is the difference structured data and entity confidence are there to describe. A site whose home page announces an Organization and whose article pages announce nothing has marked up the one page least likely to be the answer to a question. The scoring this scanner applies to a single page is set out in the methodology, and the honest reading of these figures is that a single-page score, including one of ours, describes a single page.
What to check on your own site, and what this did not measure
The check takes one command and no tools. Request your own home page the way a crawler does, without a browser, and count the a elements carrying an href that point at another page on your host. If that number is zero, nothing else on the page matters for discovery, because a crawler that does not render has arrived at a site with one document in it. If it is small, ask whether the pages you want quoted are among the ones you listed. You can see the same response this scan read on what GPTBot sees, and the declared identity this scanner arrives with is published at the bot page.
Then ask the second question, which the 143 answers above: does the page one click in carry what the home page carries? Structured data, readable prose and a title that names the thing are all properties of a template, and a site usually has several templates. Checking one of them is checking one of them.
Four limits bound everything here. No JavaScript was executed, so for any crawler that does render, a zero in this data may be a number that exists after rendering and not before; Google's own page says dynamically inserted anchors are crawlable, and the nine crawler vendors differ on whether they render at all. 214 hostnames refused this crawler with 403 and 31 never resolved, so 245 sites are simply absent, and they are not absent at random, which the eight reasons a scanner refuses a URL covers from the other direction. Only the first link was followed on each site, one hop deep, so nothing here describes a site's full link graph or how much of it is reachable. And the corpus is an editorial frame of 1,419 large organisations and platform-grouped sites, so every rate supports a statement about these hostnames and nothing wider, the same caveat attached to the crawlability study.
Nothing in these figures says a site with few internal links is doing badly by its readers. A one-page site is a one-page site, and the 55 with no anchor at all are running applications that work perfectly in a browser. What the measurement says is narrower and harder to argue with: on 21 September 2026, for a client that reads HTML and does not run scripts, 84 of 1,091 sites in this corpus were one page long, and half the navigation problem a reader worries about was solved by their platform without them knowing. Where a site sits on that is what AI visibility means in practice, and what a scan is for.
Flow: Crawler fetches home page to Any a element with href?; Any a element with href? (no) to Site is one page to this client; Any a element with href? (yes) to Any pointing at this host?; Any pointing at this host? (no) to Only off-host or fragments; Any pointing at this host? (yes) to Second page requested; Second page requested to Answers, carries nothing; Second page requested to Answers with prose and schema.
Lantad
Published .
Every measurement this blog has published about a site's readability has been taken on its home page, and so has almost every audit a reader has ever run. That is a reasonable place to start and a poor place to stop, because a home page is one document and a site is a graph. What joins the two is the set of internal links an AI crawler can follow, and whether those links exist in the bytes a crawler receives is a separate question from whether they exist on the screen.
Common questions
Do internal links matter for AI crawlers?
They are the only way a crawler reaches a page it was not given. Google's Make your links crawlable documentation, carrying Last updated 2025-12-10 UTC when read on 21 September 2026, says Google uses links for page relevancy and discovery, and that it can generally only crawl a link if it is an a HTML element with an href attribute. On 21 September 2026, 84 of the 1,091 corpus home pages that answered Lantad carried no such link to another page on their own host, and 55 of the 84 carried no anchor element at all.
How many internal links should a home page have?
There is no published threshold and this measurement does not propose one. What it gives is a distribution: across 1,091 home pages read on 21 September 2026, the median carried 48 distinct internal paths, the mean 75, and the largest 1,573. 193 carried nine or fewer. The useful test is not the count but the coverage, which is whether the pages you want quoted in an AI answer are reachable from somewhere a crawler starts.
Why does my home page show links in the browser but none in the HTML?
Because the navigation is built after the page arrives. The anchors exist in the document a browser assembles and not in the bytes the server sent, so a client that does not execute scripts sees the page without them. Google states that anchors inserted by JavaScript are crawlable provided they use the same markup, so this is not fatal for every crawler, but it makes discovery conditional on a rendering pass that no site owner can require. Among the 84 corpus home pages offering no path on 21 September 2026, 42.9 percent of the Bubble-built sites were in that group against none of the 54 Wix and Squarespace sites.
Does structured data on my home page cover the rest of my site?
Not reliably. Of the 604 corpus home pages carrying JSON-LD on 21 September 2026, 143 linked first to an internal page that carried none, which is 23.7 percent. Structured data is emitted by a template, and a site usually has several, so the home page template saying what the organisation is tells you nothing about whether the article or product template says what an article or product is.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.