BlogFindings
Blocking AI crawlers cost large publishers about 7 percent of traffic
A staggered difference in differences study of news publishers, arXiv 2512.24968, estimates roughly a 7 percent traffic decline within 6 weeks of first disallowing a generative AI crawler. Split by publisher rank, the decline sits in the top 50 and the point estimate turns slightly positive below rank 100.
The headline is about 7 percent within six weeks. The more useful part is what sits underneath that average, because the paper breaks it apart by publisher size and the pieces do not all point the same way. Lantad ran none of this. What follows is a reading of one paper, opened and quoted on 8 August 2026, and this site holds no traffic data for any publisher in its sample. The part this site can add comes after the figures, and it is the difference between a decision about crawler policy and the mechanical question of whether the decision your organisation believes it made is the one your server is actually enforcing.
In short
- Strategic Response of News Publishers to Generative AI, arXiv 2512.24968 by Hangcheng Zhao and Ron Berman, first posted 31 December 2025 and revised to a fourth version on 15 April 2026, estimates that publishers who blocked generative AI crawlers in robots.txt saw approximately a 7 percent decline in traffic within 6 weeks.
- The paper's average treatment effect on the treated is -0.074 on SimilarWeb with a 95 percent interval of [-0.141, -0.007], -0.069 on Semrush with [-0.145, 0.007] and -0.065 on Comscore with [-0.150, 0.021], so of the three panels only the SimilarWeb interval excludes zero.
- Split by Semrush rank across the top 500 news publisher domains, the paper reports -0.0691 for the top 50, -0.0343 for ranks 51 to 100 and a small positive 0.0164 for ranks 101 to 500, with all three of those intervals crossing zero.
- The paper restricts its main analysis to the pre-May 2024 period and states that its traffic data do not capture consumption taking place entirely within AI interfaces, so it does not describe what blocking costs in the era of AI Overviews.
- Lantad ran no part of this study, holds no traffic data for any publisher in its sample and no scan of any domain it names, and reports these figures from the paper as read on 8 August 2026.
| Traffic panel | Estimate (ATT) | 95 percent interval | Interval excludes zero |
|---|---|---|---|
| SimilarWeb | -0.074 | [-0.141, -0.007] | Yes |
| Semrush | -0.069 | [-0.145, 0.007] | No |
| Comscore | -0.065 | [-0.150, 0.021] | No |
What the study measured, and on how many publishers
The paper is Strategic Response of News Publishers to Generative AI by Hangcheng Zhao and Ron Berman, arXiv identifier 2512.24968, filed under General Economics with cross lists to Artificial Intelligence, Computers and Society, and Applications. It was first posted on 31 December 2025 and revised three times, reaching a fourth version on 15 April 2026, which is the version read here on 8 August 2026.
The core sample is smaller than the subject suggests. The authors start from websites that appear in their SimilarWeb dataset and also carry a matching record of employer sponsored job postings, then filter to NAICS code 513110 and to newspaper publishers. In the paper's own words, this process yields 30 URLs. Thirty is the number to hold on to when reading anything downstream, because it makes this a study of the top of American newspaper publishing rather than a sample of the web. Wider analyses extend to the top 500 news publisher domains by Comscore traffic, and that extension is where the most interesting result in the paper lives.
Blocking dates are not self reported, which matters more than it sounds. The authors identify when each publisher first disallowed a generative AI crawler by reading historical robots.txt files from the HTTP Archive, because a policy change of this kind usually arrives with no announcement attached and the file is the only public record that it happened. That is the same class of evidence behind the finding that the sites that block AI crawlers are the ones with editors. The paper states that roughly 75 percent of publishers in its sample block an OpenAI related crawler at some point, with adoption beginning as early as mid 2023 and arriving at different times for different publishers, which is exactly the staggered timing a difference in differences design needs in order to separate the effect of blocking from whatever else was happening to news traffic that year.
Traffic comes from three independent panels rather than one. SimilarWeb supplies daily domain level estimates of total visits across desktop and mobile from 1 January 2019 through 28 February 2026. Comscore supplies a desktop household browsing panel covering 2022 to 2024, aggregated to publisher week. Semrush supplies daily channel specific visits. Using three vendors with different collection methods is a deliberate robustness choice, and the results below are reported per vendor rather than pooled, which lets a reader see how much the answer depends on whose panel you believe. Historical page level HTML from the HTTP Archive, annual URL counts from the Internet Archive and monthly job postings from Revelio Labs carry the parts of the paper that are not about traffic.
The list of blocked tokens in the paper's appendix is long and spans nine vendors: gptbot, chatgpt-user and oai-searchbot for OpenAI, claudebot, claude-user, claude-searchbot and anthropic-ai for Anthropic, perplexitybot and perplexity-user, google-extended, applebot-extended, meta-externalagent and meta-externalfetcher, amazonbot, bytespider, and two Cohere tokens. Maintaining a list that size by hand, correctly spelled and kept current, is the problem behind the crawler tokens that never appear in your logs, and this site's crawler directory exists because the set keeps moving.
Flow: SimilarWeb site list to Match Revelio postings; Match Revelio postings to NAICS 513110 newspapers; NAICS 513110 newspapers to 30 URLs core sample; 30 URLs core sample (treated units) to Staggered DiD estimate; HTTP Archive robots.txt to First disallow date; First disallow date (treatment timing) to Staggered DiD estimate; Staggered DiD estimate to Traffic effect.
How large the traffic effect was, and how certain
The headline is one sentence in the paper: publishers experience approximately a 7 percent decline in traffic within 6 weeks after blocking generative AI crawlers, whether measured by SimilarWeb, Semrush or Comscore, relative to the pre-blocking period. Three panels, one direction, similar magnitude. That agreement is the strongest thing in the result, because the three vendors do not share a measurement method and would not be expected to produce a matching artefact.
The coefficients themselves are worth reading rather than the summary of them. The staggered difference in differences estimates of the average treatment effect on the treated are -0.074 for SimilarWeb with a 95 percent confidence interval of [-0.141, -0.007], -0.069 for Semrush with [-0.145, 0.007], and -0.065 for Comscore with [-0.150, 0.021]. Read those intervals rather than the point estimates. Only the SimilarWeb interval excludes zero. The Semrush and Comscore intervals both cross it, which means that on either of those panels taken alone the data are also consistent with no effect at all, and with a small positive one.
That is not a reason to discard the finding. Three estimates landing within a hundredth of each other, produced from three panels built differently, is evidence of something real, and the paper reports further specifications running the same way: a synthetic difference in differences estimate of -0.0912 on SimilarWeb, and two way fixed effects estimates of -0.0727 on SimilarWeb, -0.0404 on Semrush and -0.0647 on Comscore. It is a reason to state the finding the way the paper states it, as an approximate 7 percent with real uncertainty around it, rather than as a fact about what blocking will cost any particular site.
The mechanism the number implies is worth naming too, because it is not the one the framing invites. Blocking a crawler whose stated job is model training cannot mechanically remove a page from a search index, and the decline shows up in human browsing panels rather than in bot request counts. The plausible route runs through retrieval and referral rather than through training, which is the distinction behind the report that AI Overviews retrieved less from sites blocking Google-Extended, and behind the finding that citations reached 6.8 percent of ChatGPT prompts while the visit that followed usually landed somewhere else. The paper does not decompose its estimate into a channel, and this post is not going to do it on the paper's behalf.
Which publishers actually lost the traffic
The average is the least useful figure in the paper, and the paper says so by breaking it apart. Extending the sample to the top 500 news publisher domains and splitting them by Semrush rank produces three estimates that do not tell one story. Publishers in the top 50 show -0.0691, with an interval of [-0.145, 0.00686]. Ranks 51 to 100 show -0.0343, with [-0.141, 0.0725]. Ranks 101 to 500 show 0.0164, with [-0.0971, 0.130], which is a small positive point estimate rather than a decline.
Every one of those three intervals crosses zero, so none of the bands supports a confident claim taken on its own. The pattern across them is the finding rather than any single row: the magnitude shrinks as you move down the rank list, and the sign flips somewhere below rank 100. The Comscore split runs the same way, with declines of roughly 4.7 percent for the most heavily visited publishers and roughly 5.6 percent in the middle band, and a small positive point estimate for the lightest.
The reading that follows is uncomfortable for both sides of the usual argument. If you are a top 50 American newspaper, this paper is evidence that blocking carries a real cost. If you are the three hundredth, it is evidence of nothing measurable in either direction, and anyone quoting 7 percent at you is quoting a number estimated on a population you are not in. The asymmetry is not surprising once it is stated plainly. The publishers with the most to lose from blocking are the ones a retrieval system was most likely to reach for in the first place, and a site that was rarely cited before cannot lose citations it was not receiving.
This is the second time this blog has met the same shape in a fortnight. The sites that block AI crawlers are the ones with editors reported that blocking is concentrated among reputable news publishers rather than spread evenly across the web, and this paper reports that the measurable consequence of blocking is concentrated in the same place. Both describe a market where crawler policy is being set, and studied, at the very top, while the advice derived from it is addressed to everybody. The same caution applies to the commercial layer, where 17.0 percent of news outlets billed an AI crawler and no robots.txt said so: the publishers with the leverage to charge are a small and unrepresentative set. Anyone reading a generative engine optimization claim about crawler blocking should ask which population it was estimated on before applying it to their own site.
| Rank band | Estimate (ATT) | 95 percent interval | Direction of point estimate |
|---|---|---|---|
| Top 50 | -0.0691 | [-0.145, 0.00686] | Decline |
| 51 to 100 | -0.0343 | [-0.141, 0.0725] | Smaller decline |
| 101 to 500 | 0.0164 | [-0.0971, 0.130] | Slight increase |
What the estimate cannot tell you about your own site
The paper is explicit about its own limits, and they are load bearing rather than decorative. It states that the traffic data the authors have access to do not capture consumption that takes place entirely within AI interfaces, and that the study period precedes the wider adoption of AI integrated search products. Both halves matter. A reader who gets their answer inside a chat interface and never clicks through is invisible to SimilarWeb, to Semrush and to a Comscore browsing panel alike, so the measured decline is a decline in visits and not a decline in being read. Those two things have been drifting apart for two years.
The timing constraint is sharper still. The authors write that they restrict their main analysis to the pre-May 2024 period, because blocking adoption was concentrated in mid to late 2023, well before AI Overviews had an impact. The estimate therefore describes what blocking cost in a world where AI generated answers were not yet embedded in the dominant search results page. That world no longer exists. This blog has reported one attempt at measuring the newer one, where AI Overviews cut English Wikipedia traffic by about 15 percent, and whether the cost of blocking has risen or fallen since May 2024 is simply not a question this paper answers. Nothing here should be read as though it does.
Three further boundaries are worth stating plainly. The sample is American newspaper publishers, so nothing in it generalises to software, retail or documentation sites without a separate argument that the paper does not make. The treatment is the first disallow of a generative AI crawler, which pools a publisher who blocked one token with a publisher who blocked fifteen, so the estimate averages over policies that differ enormously in scope. And the outcome is total site traffic, which moves for many reasons inside a six week window, which is what the design exists to handle and what those confidence intervals are honestly reporting.
None of this makes the paper weak. It makes it a specific measurement of a specific thing, which is the only kind worth citing. It is also exactly the sort of qualification that disappears when a study is compressed into a headline figure, which is why this post links the paper itself rather than a summary of it, and why the methodology page for this site's own score is written to name its own boundaries in the same way. A number reported without its population and its window is not a finding, it is a slogan with a decimal point.
-
Consumption inside AI interfacesNot captured The paper states its traffic data do not capture consumption that takes place entirely within AI interfaces. -
The period after May 2024Out of scope The main analysis is restricted to the pre-May 2024 period, before AI Overviews had an impact. -
Sites that are not newsNot sampled The core sample is 30 URLs filtered to NAICS 513110 newspaper publishers, extended to the top 500 news domains. -
Which token was blockedPooled Treatment is the first disallow of a generative AI crawler, whether the file named one token or fifteen.
What a site owner can check, given the paper cannot answer it for them
The paper measures the consequence of a decision. It cannot tell you whether the decision your organisation believes it made is the one your server is enforcing, and in practice that gap is where most of the surprises live. Four checks are cheap, none of them requires a study, and all four fail silently when they fail.
The first is to read the file rather than the intention behind it. A robots.txt group is matched to a product token, and under the Robots Exclusion Protocol a request that matches no group and finds no wildcard has no rules applying to it at all. That is how a renamed crawler token leaves your robots.txt group matching nothing with no error surfacing anywhere, and the vendor rename is not a rare event. Checking a specific token against your live file is what the robots.txt tester is for, and it is a one minute job.
The second is to check what the absence of the file does. A missing robots.txt is not a neutral state, and the two common failure statuses point in opposite directions, which is the subject of robots.txt 404 and 503 are opposites. An organisation that believes it is blocking, and serves the wrong status on that file during an incident, has changed its policy without deciding to.
The third is to remember that the rule and the response are two different measurements. A disallow line is a request rather than an enforcement, and this blog has published a case where a GPTBot ban served a 200 anyway. The identity being matched is itself only an assertion, since a user agent is a claim, not an identity, and the way those layers stack is set out in two layers decide if AI can read your site. A study that dates a policy change from the file is dating the declaration, which is the right thing for it to measure and not the same thing as enforcement.
The fourth is to allow for delay. An edit does not take effect at the moment you save it, because crawlers cache the file, and when a robots.txt edit reaches a crawler sets out the window that follows. A six week measurement window in a study of this kind is not an arbitrary choice.
One disclosure about how this site treats the same decision, since it is a setting rather than a finding. Lantad's crawler registry sorts every token into three purpose classes, training, search and user_agent, and only the last two appear in the constant that decides whether a disallow costs points in the Access sub-score. Blocking a training crawler costs a site nothing in the score, by design, on the reasoning that it is a legitimate owner choice. That is a decision somebody made and not something anyone measured. This paper is a reason to keep looking at it rather than a reason to change it, because the publishers it studies began by blocking tokens documented for training, including those in OpenAI's crawler documentation, and a traffic panel and a readability score are answering different questions about the same file.
- Does the token still match a group? A renamed or newly published product token falls through to the wildcard, or to no rules at all when there is no wildcard group.
- What status does the file return when it is missing? A 404 and a 503 are read in opposite directions by a compliant crawler, so an incident can silently invert your policy.
- Does the response match the rule? A disallow in the file is a request. What the origin actually returns to that user agent is a separate measurement.
- How long until an edit is seen? Crawlers cache robots.txt, so the file you saved this morning is not the file every crawler is currently applying.
Lantad
Published .
Nearly every argument about whether to block an AI crawler is conducted without a price attached. One side says a model trained on your archive pays you nothing, the other says a crawler that cannot read you cannot cite you, and both are arguing about a quantity neither has measured. A working paper on arXiv puts a number on one half of it. It uses the staggered dates on which news publishers actually edited their robots.txt files as a natural experiment, and estimates what happened to their traffic in the weeks after.
Common questions
Does blocking AI crawlers reduce a publisher's traffic?
One causal estimate says yes for large news publishers. Strategic Response of News Publishers to Generative AI, arXiv 2512.24968, reports approximately a 7 percent decline in traffic within 6 weeks of a publisher first disallowing a generative AI crawler in robots.txt. The estimate comes from a staggered difference in differences design across three separate traffic panels, and of those three only the SimilarWeb confidence interval excludes zero. Lantad did not run this study and reports the figures from the paper as read on 8 August 2026.
Does the 7 percent figure apply to a small site?
The paper gives no support for that. When the authors extend to the top 500 news publisher domains and split by Semrush rank, the estimate is -0.0691 for the top 50, -0.0343 for ranks 51 to 100, and a slightly positive 0.0164 for ranks 101 to 500. All three of those intervals cross zero, and the direction of the point estimate flips below rank 100. The measurable cost sits at the top of the market, on a sample of American newspaper publishers.
Does the study cover the AI Overviews era?
No, and it says so. The authors restrict their main analysis to the pre-May 2024 period because blocking adoption was concentrated in mid to late 2023. They also state that their traffic data do not capture consumption taking place entirely within AI interfaces. So the estimate describes what blocking cost before AI generated answers were embedded in the dominant search results page, and it does not measure readers who never click through at all.
Does blocking a training crawler affect a Lantad score?
No. Lantad sorts crawler tokens into three purpose classes, training, search and user_agent, and only the search and user_agent classes appear in the constant that decides whether a robots.txt disallow costs points in the Access sub-score. Blocking a training crawler is treated as a deliberate owner choice that costs nothing. That is a configured setting in the repository rather than a measurement, and it is a separate question from what a traffic panel would record.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.