BlogFindings

Do cookie banners block AI crawlers? 269 of 1,093 home pages ran a consent platform, and it cost 0.44 percent of the words

Lantad requested the home page of all 1,419 hostnames in this repository's committed corpus on 21 September 2026 and read the raw bytes with no JavaScript executed. 1,093 answered HTTP 200 with HTML. 269 of them loaded one of 25 named consent platforms and 257 carried consent wording in the delivered HTML, but that wording came to 6,143 words against 1,381,167 visible words in total, and no page in the corpus gave a cookie notice more than half its text.

16 min read Lantad

This blog has already published the documentation half of the argument, in a post on how a consent gate never opens for a stateless crawler, and that post said plainly in its own takeaways that Lantad had not measured how often the gate actually costs a site its text. This is that measurement. On 21 September 2026 Lantad requested the home page of all 1,419 hostnames in this repository's two committed corpus seed files as LantadBot, followed redirects, executed no JavaScript, and read the bytes that came back. The result does not support the wall. Across the pages that answered, consent notices came to 0.44 percent of the words an AI crawler could read, and the pages carrying a consent platform returned more text than the pages carrying none.

In short

  • Lantad requested the home page of 1,419 hostnames on 21 September 2026 and read the raw bytes with no JavaScript executed. 1,093 answered HTTP 200 with HTML, and consent notices accounted for 6,143 of the 1,381,167 visible words on them, which is 0.44 percent.
  • Do cookie banners block AI crawlers on the pages that run them? Not on this corpus: the 269 pages loading a named consent platform returned a median of 1,275 visible words against 850 on the 824 pages that loaded none.
  • On 137 of the 269 pages running a consent platform, not one word of the notice appeared in the served HTML, because the platform's tag inserts the banner in the browser and a client that runs no JavaScript never reaches that step.
  • Nine of the 1,093 home pages gave a consent notice more than a tenth of their visible words and three gave it more than a quarter, the highest being famousbrands.co.za at 373 of 964 words on 21 September 2026.
  • Three pages running a consent platform returned no visible words at all, and on all three the cause was client-side rendering rather than the notice: bloomandwild.com, croatia.hr and skyscanner.net each served an empty application shell.
  • Answered HTTP 200 with HTML 1093
  • Either consent signal present 394
  • Named consent platform loaded 269
  • Consent wording present in the HTML 257
  • Platform loaded, none of its words in the HTML 137
Home pages in the 1,419 host corpus, by consent signal found in the raw bytes on 21 September 2026. 1,094 hostnames answered HTTP 200 and 1,093 of those with an HTML content type, which is the denominator for every figure in this post.

What was measured, and how a banner was counted

Each of the 1,419 hostnames received one request for its home page over HTTPS, from one network location, with redirects followed and a twenty second timeout. 31 never returned a status at all, 19 failing DNS resolution, 10 timing out and 2 failing in the client. 213 answered HTTP 403 to this crawler and 53 answered HTTP 503. 1,094 answered HTTP 200, and 1,093 of those carried an HTML content type. Those 1,093 pages are the denominator throughout, and they are not a random sample of the web: a site that refuses this scanner never reaches the test, so the surviving set leans toward sites that admit crawlers. The request discipline is the same one described on the methodology page and used for every corpus measurement on this blog.

A consent layer was counted two separate ways, because the two answer different questions. The first is whether the page loads a consent platform at all, detected by a string that the platform's own tag or content delivery network leaves in the HTML: cdn.cookielaw.org and optanon for OneTrust, consent.cookiebot.com for Cookiebot, and 23 more signatures covering 25 named platforms in total. The second is whether the words of a notice are actually in the delivered document, detected by running the same extractor the scanner uses for prose parity and then flagging text blocks that use wording only a cookie notice uses, such as we use cookies, strictly necessary, reject all and manage preferences. A page can match one test and not the other, and the gap between the two turns out to be the most interesting number here. Word counts throughout are visible words after script, style, noscript, template, svg and iframe subtrees are dropped, the same basis used when this blog counted how many home pages sent a crawler zero words.

GET /, 1,419 hosts, no JavaScript

  • GET https://{host}/ User-Agent: LantadBot/1.0 (+https://lantad.co/bot) 1,419 sent
  • HTTP 200 with an HTML content type 1,093
  • match 25 named consent platform signatures in the bytes 269 pages
  • extract visible text, flag blocks using consent wording 257 pages, 504 blocks
  • consent words against all visible words 6,143 / 1,381,167
The request and the two tests applied to every response, 21 September 2026.

Which consent platforms the corpus actually runs

OneTrust is not merely the leader here, it is half the market on this sample: 135 of the 269 pages that matched any platform matched OneTrust, which is 50.2 percent. Cookiebot follows at 28, then TrustArc at 14, Usercentrics at 13, and CookieYes and Osano at 11 each. Thirteen further platforms matched between one and five pages each. 21 pages matched more than one platform, which is why the counts sum above 269: instacart.com matched three signatures, and cnn.com, aljazeera.com, hollywoodreporter.com and newegg.com each matched two. Those are usually a migration that was never finished rather than a deliberate pairing, and each one is a second tag loading on every page view.

The rate varies a great deal by the kind of site. Among travel hostnames 37 of 79 loaded a platform, the highest of any stratum, followed by SaaS at 52 of 119 and ecommerce at 28 of 66. At the other end, 6 of 87 government pages matched, 3 of 45 single page app startups, and none at all of the 54 Wix and Squarespace sites. That last figure needs care rather than a conclusion: those builders ship their own consent notice as part of the platform and inject it in the browser, so the absence of a third-party signature is not the absence of a banner. It is a limit of what a signature test can see, which is the same limit that applies whenever a scan reads only what the server sent. A reader who wants to see the delivered bytes for a single page can run what GPTBot sees against it, and the same discipline of reading only what arrived is what produced this blog's count of how many home pages offered a crawler no internal path.

Consent platformHome pagesShare of the 269
OneTrust13550.2%
Cookiebot2810.4%
TrustArc145.2%
Usercentrics134.8%
CookieYes114.1%
Osano114.1%
iubenda83.0%
Quantcast72.6%
HubSpot banner72.6%
Complianz72.6%
cookie-law-info plugin72.6%
Didomi62.2%
Consent platforms identified by a string the platform's own tag or CDN leaves in the home page HTML, 21 September 2026. 269 pages matched at least one and 21 matched more than one, so the column sums above 269. Thirteen further platforms matched between one and five pages each.

Where the notice did take words, and how few pages that was

Nine pages of 1,093 gave a consent notice more than a tenth of their visible words, and three gave it more than a quarter. famousbrands.co.za is the worst case in the corpus at 373 of 964 words, or 38.7 percent, a full cookie table rendered server-side into the document. plannedparenthood.org follows at 310 of 960 words and cshl.edu at 359 of 1,178. These are the pages where the notice is not injected by a script but written into the template, so a crawler receives the cookie policy and the page content as one undifferentiated run of text.

Whether that is a problem depends on where the words land, and the extractor answers that too. On 56 of the 1,093 pages the consent wording fell inside blocks classified as main content rather than boilerplate, which is the case that matters: text in a navigation or footer region is already discounted, and this blog has measured how much of a page that region accounts for, finding that 1,395 of 2,729 text blocks were navigation. interactivebrokers.com is the clearest example, with all 586 of its consent words landing in main. The practical cost is dilution rather than blocking. A passage about cookie categories sitting inside the main region competes with the page's actual subject for whatever the extracting system weighs, in the same way that container choices change what survives extraction: deleting one element cost a captured page 1,665 of 13,615 words. None of that is a claim that an answer engine ranks these pages lower, which is not something this measurement can see. What it can say is that on 1,084 of 1,093 pages the notice took under a tenth of the text, so for almost every site the cookie banner is not where the structured data and prose budget is being lost.

  • famousbrands.co.za 38.7% 373 of 964 words
  • plannedparenthood.org 32.3% 310 of 960 words
  • cshl.edu 30.5% 359 of 1,178 words
  • entelect.co.za 20.8% 212 of 1,017 words
  • blacksmith.sh 15.3% 159 of 1,041 words
  • interactivebrokers.com 14.9% 586 of 3,921 words
The six home pages where a consent notice took the largest share of the visible words, 21 September 2026. No page in the corpus exceeded half, and the corpus median among pages carrying any notice text was six words.

Three pages sent no words at all, and the banner was not the reason

Three of the 269 pages running a consent platform returned zero visible words to this crawler, and it would be easy to file them as the wall working exactly as advertised. They are not that, and checking why is the difference between a measurement and a headline. bloomandwild.com runs Cookiebot and returns an Angular application shell whose body holds eighteen script elements and no prose. croatia.hr runs Cookie Control and returns a body containing a noscript element that says in plain words that the application does not work properly without JavaScript. skyscanner.net runs Ketch and returns a loading spinner inside an empty application root. In all three cases the consent platform is present and irrelevant: the page withheld its text from a non-rendering client because it is built to assemble itself in the browser, which is the failure this blog measured when JavaScript supplied 7.6 percent of the prose across 271 pages. The fix for these three is a rendering decision of the kind covered in the per-stack guide for Next.js, not a consent decision.

One page did visibly react to the missing cookie, and it is instructive that it is only one. nature.com redirected the request to its own home page carrying the query string error=cookies_not_supported, which is a site explicitly noticing that the client kept no state. It then served 1,433 visible words including its full run of article headlines, of which 5 words were consent wording. The site detected the stateless client, announced the detection in the URL, and delivered the page anyway. That is the whole finding in miniature. Elsewhere the missing text has an entirely different cause, as when text inside shadow DOM reaches the browser and not the extractor, and attributing any of those absences to the cookie notice would be wrong.

  • bloomandwild.com 0 words Cookiebot present. The body is an Angular application shell carrying eighteen script elements and no prose, so the cause is client-side rendering.
  • croatia.hr 0 words Cookie Control present. The body carries a noscript element stating the application does not work properly without JavaScript.
  • skyscanner.net 0 words Ketch present. The body holds a loading spinner and an empty application root, with the page assembled in the browser.
  • nature.com 1,433 words OneTrust present. Redirected to a URL carrying error=cookies_not_supported and then served the full home page, 5 words of it consent wording.
The three corpus home pages running a consent platform that returned no visible words on 21 September 2026, and the one page that reacted to the stateless request.

What this measurement does not cover

Four limits, stated because the figures are worth less without them. First, no JavaScript was executed, so every count describes what a non-rendering client received. A crawler that renders would meet the banner as a visual overlay, and this measurement says nothing about whether such an overlay affects what an agent extracts or clicks. Second, the platform test is a signature match against 25 named products, so a hand-built banner or a builder's bundled notice is invisible to it. That cuts both ways: 125 pages carried consent wording with no platform signature matched at all, which is why the two tests are reported separately rather than merged into one number. Third, 31 hostnames never resolved and 213 refused this crawler with HTTP 403, so 326 sites are missing from every figure and they are not missing at random. Fourth, only the home page of each site was read, so nothing here describes an article template, a paywall, or a regional variant that behaves differently.

The check worth running on your own site is short and needs no tool. Fetch your own home page the way a crawler does, with no browser and no cookie jar, and read what comes back: if your article text is in those bytes, the banner is not your problem, whatever the advice says. If it is not, the cause will almost certainly be rendering rather than consent, which is a different repair with a different cost. For the neighbouring question of which crawlers your rules actually admit, the robots.txt tester evaluates a file per crawler token rather than in general. Everything in this post came from one request per host on one day from one network location, and the underlying corpus and method are the same ones used across this blog's research. A figure here supports a statement about these 1,419 hostnames and nothing wider.

  • Consent notices took 0.44 percent of visible words 6,143 words of 1,381,167 across the 1,093 home pages that answered HTTP 200 with HTML.
  • Pages with a consent platform returned more text Median 1,275 visible words against 850 without one, which describes who buys a platform rather than what it does.
  • Rendered behaviour was not measured No JavaScript was executed, so nothing here describes what a banner overlay does to an agent that renders the page.
  • Builder-bundled notices are invisible to the test None of the 54 Wix and Squarespace pages matched a third-party signature, which is a limit of signature matching rather than an absence of banners.
  • 326 hostnames are absent from every figure 31 never resolved and 213 answered HTTP 403 to this crawler, so the surviving corpus leans toward sites that admit crawlers.
What this measurement establishes and what it leaves open, 21 September 2026.

Written by

Lantad

Published .

Almost every commercial website in Europe opens with a question about cookies, and there is a widely repeated claim about what that question does to a machine reader. The claim is that the banner is a wall. A crawler arrives, meets the notice, and never reaches the article underneath, so the site is invisible to the systems that answer questions about it. The claim is plausible, it is cheap to repeat, and it is almost never accompanied by a count.

Common questions

Do cookie banners stop ChatGPT or other AI crawlers reading a page?

Not on this corpus. Across 1,093 home pages measured on 21 September 2026, consent notices accounted for 0.44 percent of the visible words, and on 137 of the 269 pages running a consent platform the banner's words were not in the delivered HTML at all because the platform builds the notice in the browser. No page in the corpus gave a cookie notice more than half its text.

Why did pages with a consent platform return more text, not less?

Because a consent platform is a purchase, and the organisations that buy one tend to be the ones that publish a lot. The 269 pages running a platform returned a median of 1,275 visible words against 850 on the 824 that ran none. That is a fact about which sites buy consent software, not evidence that the software helps, and this measurement cannot separate the two.

Which consent platform is most common?

OneTrust, on 135 of the 269 pages that matched any of the 25 named platforms on 21 September 2026, which is 50.2 percent. Cookiebot was second at 28 pages and TrustArc third at 14. 21 pages matched more than one platform signature, usually the residue of a migration.

If my home page shows a crawler no text, what is the likely cause?

Client-side rendering, not the cookie banner. All three corpus pages that ran a consent platform and returned zero visible words were application shells that assemble themselves in the browser, and the one site that explicitly detected the stateless request, nature.com, still served 1,433 visible words. Fetch your own page with no browser and read the bytes before blaming the notice.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.