BlogFindings
Do AI crawlers follow pagination? 116 of 466 archive pages declared a next page, and 24 offered only a button
Lantad requested robots.txt and then the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 3 October 2026 as LantadBot, followed each home page's own link to a blog, news or newsroom index, and read that index with no JavaScript executed. 466 index pages answered HTTP 200 with HTML. 116 of them declared a second page in markup a crawler can follow, and every one of those 116 URLs returned a working page. 24 others put the next tranche behind a load more control and declared no followable path to it at all.
That is Google describing Googlebot, not OpenAI or Anthropic describing theirs, and the distinction matters enough that this post keeps it. What Lantad can measure is the half that sits on the site: whether a second page exists in the delivered HTML at all. This run asked all 1,419 hostnames in the two committed corpus seed files for robots.txt, read every permitted home page as an AI crawler would, followed each home page's own link to a blog, news or newsroom index, and read that index the same way. 466 indexes answered. 116 of them offered a way forward.
In short
- Do AI crawlers follow pagination is really two questions, and only the second is about the crawler: a crawler can follow a next page only where the delivered HTML contains one, and on 3 October 2026 Lantad found that markup on 116 of 466 archive pages reached from the corpus home pages.
- Every one of those 116 declared next-page URLs answered HTTP 200 with HTML when Lantad requested it the same day, and all 116 returned a body of a different byte length from page one, so where the path exists it works.
- 24 of the 466 archive pages carried a load more or show more control and no followable next page anywhere in the markup, among them atlassian.com, cnbc.com, forbes.com, auth0.com and jpmorganchase.com, each reconfirmed by a second request the same day.
- The publishing platform predicts the outcome better than the industry does: 19 of the 26 archives in the corpus WordPress stratum declared a followable next page against 0 of the 21 in the Webflow stratum, 1 of 15 in the Framer stratum and 2 of 27 in the single page app stratum.
- Google's own pagination documentation, carrying Last updated 2025-12-10 UTC, states that its crawlers do not click buttons and do not generally trigger JavaScript functions requiring user actions, and separately that it no longer uses link rel=next, which 36 of the 466 pages still declared.
| Stage | Sites | What happened |
|---|---|---|
| Hostnames asked | 1,419 | The committed corpus, an editorial frame rather than a random draw |
| robots.txt closed the root | 98 | 13 by an explicit rule in a parsed file, 85 assumed from a 503 or no answer |
| Home page refused or failed | 245 | 216 answered 403, the rest a mix of 202, 429, 404 and 503 |
| Home page answered 200 | 1,076 | Of which one sent a content type that was not HTML |
| Home page readable | 1,075 | HTTP 200 with an HTML content type |
| Linked no archive-shaped page | 598 | No /blog, /news, /insights or similar index in the home page markup |
| Linked an archive page | 473 | The home page named an index a crawler could reach |
| Archive page readable | 466 | The denominator for every figure below; 7 answered 403 |
| Declared a followable next page | 116 | 88 by a numbered link, 28 by rel=next alone |
| Offered a button and nothing else | 24 | A load more control with no followable path |
Do AI crawlers follow pagination, and what would following it mean?
Three patterns put a long list in front of a reader, and Google's pagination page names all three: pagination, where numbered links move between pages that each have their own address; load more, where a button extends the list already on screen; and infinite scroll, where reaching the bottom fetches more. Only the first produces a second URL in the HTML. The other two are events, and an event needs something to fire it.
This is where the HTML specification settles the argument rather than any crawler vendor. The WHATWG standard, at html.spec.whatwg.org/multipage/text-level-semantics.html, says that if the a element has no href attribute then the element represents a placeholder for where a link might otherwise have been placed, consisting of just the element's contents. A placeholder is not a link. It is markup that looks like one to a person and is nothing to a parser, which is the same gap that makes a link named only by an aria-label hard to use as a path.
The second mechanism is the rel attribute. MDN's reference for rel defines next as indicating that the current document is part of a series and the next document in the series is the referenced one, which is exactly the relationship an archive has. It is also the mechanism Google has retired: the same pagination page states that Google no longer uses these tags, although the links may still be used by other search engines. That makes rel=next a real signal with an uncertain audience rather than a dead one, and it is worth counting separately for that reason.
None of this tells you what GPTBot or ClaudeBot do, and nothing in this post claims otherwise. Lantad has measured that two of nine AI crawlers render JavaScript at all, which bears on the question without answering it, and the vendor documentation behind how to get cited by Perplexity says nothing about clicking. The measurable half is the markup, so that is the half measured here.
Sample Illustrative, not a measurement of any real site.
Flow: Archive page one (href present) to Numbered link or rel=next; Archive page one (click required) to Load more button; Archive page one (scroll required) to Infinite scroll; Numbered link or rel=next to Second URL in the HTML; Load more button to Nothing a parser can follow; Infinite scroll to Nothing a parser can follow.
How 466 archive pages were found, and what was discarded on the way
A home page is the wrong place to look for pagination, because a home page is not usually a list. The archive is, so the crawl had to reach one the way a crawler would: from a link on the home page, not from a guessed path.
Every one of the 1,419 hostnames was asked for robots.txt first and the result evaluated for the LantadBot token at the site root with this scanner's own parser, the same evaluation behind the robots.txt tester. 1,109 files answered HTTP 200. 13 of the files that parsed as plain text carried a rule disallowing this crawler at the root, and a further 85 were treated as closed because the request returned a 503 or never answered, which is the conservative reading this scanner applies rather than a refusal anyone wrote. That left 1,321 hosts to ask for a home page, of which 1,075 answered HTTP 200 with an HTML content type. 216 answered 403, which is bot refusal rather than anything to do with pagination and is its own subject, measured separately when 103 of 1,089 sites served an unknown bot and refused GPTBot.
Each readable home page was then re-read and its own links scanned for a same-origin path whose first segment named an index: blog, news, article, insights, stories, press, newsroom, resources, updates, posts, publications, media, library, research, magazine, journal or events, with no query string and no fragment. 598 of the 1,071 home pages that answered the second read carried no such link, which is an ordinary result rather than a fault: a product site with no blog has no archive to paginate. 473 did carry one, and 466 of those pages answered HTTP 200 with HTML. Seven answered 403.
226 of the 466 were a /blog, 77 a /news, 31 /resources, 25 /events and 21 a /newsroom. Those 466 pages are the denominator for everything below, and the whole set was fetched a second time the same day, independently, to check that nothing here is a one-off. All 466 answered again, and the classification was identical on both passes: the same 116, the same 27, the same 24. The methodology page describes the general shape of this kind of capture.
LantadBot/1.0, redirects followed, no JavaScript
- GET /robots.txt 200 text/plain, LantadBot allowed at /
- GET / 200 text/html
- scan home page links for an index path found /blog
- GET /blog 200 text/html
- scan for a href matching /page/N, ?page=N, rel=next found /blog/page/2
- GET /blog/page/2 200 text/html, different byte length
116 of 466 declared a next page, and all 116 of them worked
116 of the 466 archive pages, a quarter of them, put a second page into the HTML in a form a parser can act on. 88 did it with a numbered link, matching a path segment such as /page/2 or a query parameter such as ?page=2 or ?paged=2 on the same origin. 36 declared a link rel=next in the head, 37 declared an a rel=next in the body, and 28 of the 116 relied on one of those two alone with no numbered link anywhere. 44 did the reverse, numbering the pages and declaring no rel=next at all. Not one of the 466 paginated by an offset parameter alone: the scan looked for offset, start and skip as well as page numbers, and every archive that declared a numbered link used a page index.
The more useful finding is what happened when those URLs were requested. All 116 were fetched the same day, taking the lowest page number each page declared, and all 116 answered HTTP 200 with an HTML content type. Every one returned a body of a different byte length from the page that linked it, so none of them was a silent redirect back to page one and none was an empty shell. There were no soft failures in this set, which is a different result from the 125 of 1,094 sites that answered a URL that does not exist with HTTP 200 and a genuinely good one: where sites bother to declare pagination, the declaration is honest.
The depth varies more than the presence does. The deepest page number declared on a single archive was 1,883, on nowhabersham.com, followed by gov.scot at 925, ny.gov at 579 and utoronto.ca at 485. The median deepest number across the 88 numbered archives was 13, and the shallowest was 2. A median archive therefore exposes something like thirteen pages of its own history to anything that reads HTML, which is a far larger surface than the single page a reader sees, and it is the surface that decides what an answer engine has available to cite. The counting of what is on each page is a separate question, covered by 985 links found against 200,712 URLs declared.
24 archives put the second page behind a button and nowhere else
27 of the 466 archive pages carried a control asking for more of the list, counted only where the visible text began with load more or show more, or named the thing being extended, as in more articles or load more stories. Three of the 27 also declared a followable next page, so a crawler reading those three loses nothing. On the other 24 the control was the only way forward in the markup.
The elements were ordinary. Those 27 pages carried 38 such controls between them: 18 read load more, 11 read show more, and the remaining nine named the thing being extended, among them load more posts, load more stories and show more news sections. 30 of the 38 were a button element, 7 were an a element with no href at all, which is exactly the placeholder case the HTML specification describes, and one was a div carrying role=button. None of the three forms leaves a URL behind.
The 24 are not small sites. They include atlassian.com and auth0.com and retool.com and airtable.com on their blogs, cnbc.com and forbes.com and aljazeera.com and yle.fi on their news and media indexes, jpmorganchase.com and metlife.com in finance, and clevelandclinic.org on its research index. Five sit in the corpus SaaS stratum and six in the news stratum. Each of the 24 was re-fetched the same day and every one of them presented the same control and the same absence of a next page.
Two honest qualifications belong here. The first is that these archives are not invisible: the median button-only page still exposed 58 same-origin paths of its own, against 65 on the pages that paginate, so page one is read either way. What is unreachable by this route is everything after page one. The second is that 20 of the 24 declare at least one Sitemap line in their robots.txt, and a sitemap is a separate discovery path that may well carry the deeper items. Whether it does on these 24 was not measured in this run. 839 of 1,016 robots.txt files declare a sitemap and 44 of 871 declared sitemaps delivered no file, so neither the presence nor the usefulness of one should be assumed from the line alone.
| Site | Archive | Stratum | Control text |
|---|---|---|---|
| atlassian.com | /blog | SaaS | Load More |
| auth0.com | /blog | SaaS | Load more (a element, no href) |
| airtable.com | /articles | SaaS | Load more |
| retool.com | /blog | SaaS | Show more |
| cnbc.com | /media | News | Load More |
| forbes.com | /news | News | More Articles |
| aljazeera.com | /news | News | Show more news sections |
| texastribune.org | /events | News | Load more posts |
| jpmorganchase.com | /ir/news | Finance | Load more |
| metlife.com | /stories | Finance | Show More (div, role=button) |
| clevelandclinic.org | /research | Healthcare | Show More (a element, no href) |
| wisesystems.com | /blog | SaaS marketing | Load more stories |
The platform decides this more reliably than the industry does
Splitting the 466 archives by the stratum each host sits in produces a sharper split than splitting them by sector, and the direction is consistent enough to be worth stating plainly.
19 of the 26 archives in the corpus WordPress stratum declared a followable next page, the highest rate anywhere in the set and roughly three times the overall quarter. That is not a quality judgement about those sites. It is the default: numbered archive pagination has been in the stock WordPress templates for years, so the sites get it without deciding to. A default is not a strategy, and the same stratum is careless elsewhere: 16 of 24 WordPress sites named no AI crawler in robots.txt.
The modern builders go the other way. 0 of the 21 archives in the Webflow stratum declared a followable next page. 1 of 15 did in the Framer stratum, and 2 of 27 in the single page app stratum. Those three strata contribute 63 archives and 3 followable next pages between them, against 19 from 26 WordPress sites. A collection list with a load more interaction is the path of least resistance in those tools, and the result is an archive whose history exists only after a click. This is the same shape as 20 of 40 Webflow sites holding no robots.txt rule and the inverse of the markdown surface 35 of 36 Framer sites served, which is a reminder that a platform can be generous on one axis and closed on another.
By sector the spread is narrower and less interpretable. News archives declared a followable next page on 4 of 24 and supplied 6 of the 24 button-only cases, more than any other stratum, which is an awkward result for the one category whose entire asset is a dated back catalogue, and it sits oddly beside the finding that 42 of 77 robots.txt files blocking a citation crawler were news sites. Education managed 11 of 50 and government 6 of 24. For anyone auditing their own stack, the per-platform notes at fix for Next.js and fix for React cover the rendering half of the same problem.
| Stratum | Archives read | Followable next page | Button only |
|---|---|---|---|
| WordPress | 26 | 19 | 0 |
| Wix and Squarespace | 16 | 8 | 0 |
| Travel | 18 | 8 | 0 |
| SaaS | 94 | 26 | 5 |
| Education | 50 | 11 | 0 |
| Healthcare | 36 | 8 | 2 |
| SaaS marketing | 34 | 8 | 5 |
| Government | 24 | 6 | 0 |
| Static docs | 23 | 5 | 0 |
| Media local | 17 | 5 | 1 |
| News | 24 | 4 | 6 |
| Finance | 23 | 2 | 2 |
| Single page apps | 27 | 2 | 1 |
| Framer | 15 | 1 | 1 |
| Webflow | 21 | 0 | 1 |
What this does not measure
This is a census of markup, not of crawler behaviour, and the gap between those two is where most claims about AI visibility go wrong. No crawl by GPTBot, ClaudeBot or PerplexityBot was observed, no server logs were held, and no answer engine was asked to cite any of these archives. Nothing here supports a statement that any engine did or did not reach page two of anything. What it supports is narrower and still useful: on 24 of 466 archives, the markup offers no page two to reach.
Four specific limits are worth naming. One archive page per host was read, chosen as the first index-shaped link on the home page, so a site whose /blog paginates badly while its /news paginates well is recorded by whichever came first. The 598 home pages carrying no index-shaped link were dropped rather than searched harder, so the 466 is a floor on how many archives the corpus holds. Nothing behind any load more control was retrieved, because retrieving it needs the JavaScript execution this scan deliberately withholds, so the volume of content hidden on those 24 sites is unknown and is not estimated here. And the corpus is an editorial sampling frame built for platform and industry coverage rather than a random draw, so every rate above describes these 1,419 hostnames and nothing wider, a constraint that applies equally to the crawlability study.
One more thing is outside this measurement and inside the subject. Google retired link rel=next as an indexing signal and said so, yet 36 of these 466 pages still declare it. That is not a defect and this post does not score it as one. It is a reasonable hedge: the attribute costs nothing, MDN still documents it as a series relationship, and the engines that matter for AI visibility have published nothing either way about reading it. Checking what your own archive hands a crawler takes one request, and what GPTBot sees will show you the delivered bytes.
-
Markup presentMeasured 116 of 466 archives declared a followable next page, reconfirmed on a second pass the same day -
Next URL resolvesMeasured 116 of 116 answered HTTP 200 with HTML and differed in byte length from page one -
Control presentMeasured 27 archives carried a load more control, 24 of them with no followable alternative -
Content behind the buttonNot measured Retrieving it requires executing JavaScript, which this scan withholds by design -
Whether an AI crawler reached page twoNot measured No crawler was observed and no server logs were held -
Whether sitemaps cover the gapNot measured 20 of the 24 declare a Sitemap line; what it contains was not read this run
Lantad
Published .
A crawler that executes no JavaScript and clicks nothing reads the bytes it is given and follows the links inside them. That is the entire mechanism, and it sets a ceiling on how much of a site an answer engine can ever quote. So the question do AI crawlers follow pagination turns out to be two questions stacked on top of each other, and only the second one is about the crawler at all. The first is whether the page hands it a next page to follow. Google states the rule for its own fetchers without hedging: Google can generally only crawl a link that is an a element with an href attribute, on a page carrying Last updated 2025-12-10 UTC, which adds that it cannot reliably extract URLs from a elements that have no href. Its pagination and incremental page loading guidance, carrying the same date, is blunter: Google's crawlers do not click buttons and generally do not trigger JavaScript functions that require user actions to update the page.
Common questions
Do AI crawlers follow pagination?
A crawler can only follow a next page that exists in the delivered HTML as an a element with an href, or as a rel=next link. Google's pagination documentation, carrying Last updated 2025-12-10 UTC, states that its crawlers do not click buttons and do not generally trigger JavaScript functions requiring user actions. Lantad measured the site half of this on 3 October 2026: 116 of 466 archive pages reached from the corpus home pages declared such a path, and 24 others offered only a button.
Is a load more button bad for AI visibility?
It is only a problem where nothing else declares the next page. Three of the 27 archives Lantad found carrying a load more control on 3 October 2026 also declared a followable next page, so a crawler reading those three loses nothing. On the other 24 the control was the only route past page one in the markup. The usual fix is to keep the button and add numbered links behind it, which is what Google's pagination guidance asks for.
Does link rel=next still do anything?
Google says it no longer uses link rel=next or rel=prev, although it notes the links may still be used by other search engines. 36 of the 466 archive pages Lantad read on 3 October 2026 declared one in the head and 37 declared an a rel=next in the body, with 28 pages relying on one of those alone and carrying no numbered link. No AI crawler vendor has published whether it reads the attribute.
Which platforms paginate in a way a crawler can follow?
In this corpus the stock templates did better than the modern builders. 19 of the 26 WordPress archives Lantad read on 3 October 2026 declared a followable next page, against 0 of 21 in the Webflow stratum, 1 of 15 in the Framer stratum and 2 of 27 in the single page app stratum. Numbered archive pagination is a WordPress default, while a collection list with a load more interaction is the default in the builders.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.