Blog / 17.0 percent of news outlets billed an AI crawler, and no robots.txt said so

17.0 percent of news outlets billed an AI crawler, and no robots.txt said so

A probe of 185 media properties published on 4 August 2026 reports that 17.0 percent of the outlets giving a clean baseline answer at least one AI vendor with HTTP 402 Payment Required, that the gate is selective by vendor, and that none of it is declared in robots.txt.

In short

  • Kin Lane published a probe of 185 media properties at API Evangelist on 4 August 2026, reporting that of the 153 outlets giving a clean human baseline, 17.0 percent gate at least one AI vendor behind HTTP 402 Payment Required, 11.8 percent gate selectively, 30.7 percent block agents outright with no price offered, and 52.3 percent remain open to every vendor.
  • The same probe reports Anthropic billed by 15.0 percent of those outlets against OpenAI's 7.8 percent, with Google reaching 79.1 percent of outlets unpaid and Apple 78.4 percent, and six outlets letting OpenAI in free while billing Anthropic against two doing the reverse.
  • Gating in that panel clusters by owner and by category rather than by masthead: entertainment and lifestyle outlets gate at 45.8 percent, science and health at 5.9 percent, and the technology trade press at 0.0 percent.
  • Differential serving, meaning a site handing an agent a different document from the one it hands a person, appeared on 1.8 percent of the 436 sites in the first pass, with Time the only publisher in the panel serving advertising payload to agents.
  • Lantad did not run this probe, operates no publisher panel, and publishes no figure for how many sites charge for crawler access, which is a limit of what any robots.txt reader can see rather than a gap in this particular scanner.

Everything this scanner says about a page comes out of one exchange. A named client asks for a URL, the server answers with a status code and some headers, and either the words arrive or they do not. That exchange is where an AI crawler succeeds or fails, and it is the layer our methodology is built on top of.

A probe published on 4 August 2026 pointed the same exchange at a question this site has never asked. Kin Lane, writing at API Evangelist, sent requests to 185 media properties under a range of named AI vendor user agents and recorded not whether the page came back, but what came back instead. On a substantial minority of them the answer was HTTP 402 Payment Required, carrying a machine-readable body naming a toll collector or a licensing address. Lantad ran none of this, operates no publisher panel, and holds no figure of its own for how many sites charge for access. What follows reports Lane's numbers with his method and his exclusions attached, because both are unusually well documented, and then takes up the part that is genuinely this site's subject: none of what he found is written in any of those publishers' robots.txt files, and no amount of reading robots.txt would have surfaced it. The article is at https://apievangelist.com/2026/08/04/the-ai-licensing-market-is-visible-in-http-status-codes/ and is not linked here because the host is not on this site's outbound register.

  • Anthropic billed by 15.0% Roughly double OpenAI's exposure in the same panel
  • OpenAI billed by 7.8% Reaches 66.7 percent of outlets without paying
  • Apple billed by 5.9% Reaches 78.4 percent unpaid
  • Google billed by 3.9% Reaches 79.1 percent unpaid, the widest free reach in the panel
Share of the 153 clean-baseline outlets that answer each vendor with a bill rather than a page. Source: Kin Lane, API Evangelist, 185 properties probed, 153 clean baseline, published 4 August 2026. Not a Lantad measurement.

What the probe measured, and the 17.3 percent it threw away

The shape of the work matters more than any single number in it, so it is worth setting out before the findings. Lane started with 436 sites while looking for something else entirely, then widened to 185 media properties across eight categories once the interesting result turned up. Of those 185, he reports that 153 produced a clean human baseline. The other 17.3 percent were discarded, and the reason he gives for discarding them is the most useful sentence in the piece: those outlets bot-defend every client, including a plain desktop browser user agent, so nothing can be learned about their agent policy from outside. He names Reuters as one, answering a desktop Chrome user agent with a 401.

That exclusion is the difference between a figure and a headline. An outlet that refuses everybody is not an outlet that has taken a position on AI crawlers, and counting it as one would have inflated the blocking rate by about a third on his own estimate. The choice to publish the exclusion rate next to every rate that depends on it is the practice this site tries to hold itself to when it withholds a grade rather than guessing one, and it is rare enough in crawler research to be worth naming.

The tooling detail is equally load bearing. Lane reports running two independent HTTP clients, curl and Python, and refusing to report a result when the two disagree. He explains why in the same post: his first pass at the earlier Time teardown was probed with curl alone and produced a tidy story about which crawlers were blocked that turned out to be an artefact of curl's TLS fingerprint, because the identical user agent sent from Python got a different answer. That is not a footnote. It is the same failure mode this site wrote up when nine coding agents arrived at an origin as curl and axios, and the same one behind the finding that headless Chromium was refused on 15.2 percent of the top 10,000 sites: the client string is one of several things an edge inspects, and a probe that varies only the user agent is measuring less than it thinks. Lane also published two other errors he corrected on the way, having first labelled the 402 responses cloaking, which would have been a false accusation, and having first collapsed blocked and billed into one bucket, which inverted a finding. A methods section that lists what the author got wrong is a stronger signal than one that does not.

StageCount or rateWhat it means
First pass, all sites436Original probe, looking for differential serving
Media panel185Widened across eight categories
Clean human baseline153The denominator for every rate below
Excluded, defends all clients17.3%Refuses a plain browser too, so agent policy is unmeasurable
Open to every vendor52.3%No toll, no block
Blocked outright, no price30.7%Refusal rather than a bill
Gated behind a 40217.0%At least one vendor billed
Gated selectively11.8%Some vendors in free, others billed
The panel as reported by Kin Lane, API Evangelist, 4 August 2026. Percentages are of the 153 clean-baseline outlets unless stated.

Which AI companies get billed, and which walk in free

The reason a per-vendor breakdown exists at all is that the gate is not a gate. It is a list. Lane reports The Atlantic answering OpenAI's crawler with a 200 and Anthropic's with a 402, the Associated Press doing the same, the San Francisco Chronicle letting OpenAI and Perplexity through while billing Anthropic, and Forbes billing everybody. A single site therefore holds several different answers at once, selected by which name the requesting client puts in its user agent header.

Across the 153 outlets the reported reach figures are Google at 79.1 percent unpaid and billed by 3.9 percent, Apple at 78.4 and 5.9, OpenAI at 66.7 and 7.8, and Anthropic at 64.7 and 15.0. Head to head, six outlets let OpenAI in free while billing Anthropic, and two do the reverse. Lane is careful about what that does and does not establish, and the caution deserves repeating rather than paraphrasing away: a 402 means no current paid access, not refusal, so some of those responses are live negotiations and some are lapsed renewals. It is a snapshot of which deals are papered, not a ranking of which companies publishers like.

There is a structural point underneath the table that this site has written about from the other direction. A per-vendor answer only exists because the vendors publish distinct tokens to be told apart by, which is why six of the nine AI vendors we track publish exactly one crawler token matters commercially and not just technically. A publisher cannot bill a company it cannot name in an edge rule, and it cannot name one that does not identify itself consistently. The same property cuts the other way and is the reason none of this is enforcement: a user agent string is a claim rather than an identity, so the toll applies to clients that announce themselves honestly and to nobody else. Anyone probing this way, Lane included, is measuring how an origin treats a claimed name, not how it treats a verified one.

VendorReaches unpaidBilled byNotes from the report
Google79.1%3.9%Widest free reach in the panel
Apple78.4%5.9%Second widest free reach
OpenAI66.7%7.8%Six outlets admit it free while billing Anthropic
Anthropic64.7%15.0%Roughly double OpenAI's billed exposure
Vendor reach across the 153 clean-baseline outlets, as reported by Kin Lane, API Evangelist, 4 August 2026. Reported figures, not measured by Lantad.

The postures cluster by owner, not by newsroom

The expectation going in, Lane says, was that individual newsrooms were making these calls. The data he reports says otherwise, and the pattern is the sort of thing that only becomes visible once you probe many properties in one sweep rather than auditing one site at a time.

Every People Inc property in the panel runs an identical first-party toll: People, Investopedia, Allrecipes, Serious Eats, Travel and Leisure and Verywell Health all block OpenAI, Google and Amazon outright while billing the other eight vendors. Six mastheads, one decision, and Lane reports it confirmed by Investopedia's own 402 response naming People Inc's content-licensing address. Hearst splits by division instead, its newspapers billing Anthropic and nobody else while Esquire on the magazine side bills only CommonCrawl. Penske Media, which acquired Vox Media in June, is running two policies at once: Deadline, Variety and Rolling Stone bill nine vendors through TollBit, while the newly acquired Verge, Vox and Polygon run no toll at all. Of the gated outlets, fourteen collect through TollBit and twelve run a first-party toll.

The category breakdown is the part worth sitting with. Entertainment and lifestyle publishers are the most heavily gated at 45.8 percent, science and health sit at 5.9 percent, and the technology trade press gates at 0.0 percent, not one outlet. The publications explaining this fight to everyone else are, on this measurement, giving their work to every model that asks. It is a different cut of the same territory as the arXiv study finding that the sites which block AI crawlers are the ones with editors, and the two do not contradict each other so much as measure different layers: that study read robots.txt files, and this one read status codes at the edge. A site can be permissive in the file and expensive at the door, which is precisely the divergence behind 234 of 592 sites that ban GPTBot in robots.txt serving it a 200 anyway, running here in the opposite direction.

  • Entertainment and lifestyle 45.8% The most heavily gated category reported
  • All outlets in the panel 17.0% The headline rate
  • Science and health 5.9% Near the bottom of the reported range
  • Technology trade press 0.0% Not one outlet in the category gates a vendor
Share of outlets gating at least one AI vendor behind a 402, by category. Source: Kin Lane, API Evangelist, 4 August 2026, 153 clean-baseline outlets across eight categories. Three of the eight categories are quoted in the report.

The agent advertising story is one publisher

The probe began as a search for something else, and the negative result it produced corrects an expectation this blog helped carry. On 1 August 2026 we wrote up trade reporting that Time was selling sponsored blocks inside the markdown AI crawlers read, and framed it as the first instance of a page that could differ from its human twin by commercial design. That framing was accurate about Time and said nothing about how common the practice was, because nothing had counted it.

Something has now. Lane reports differential serving, meaning a site handing an agent a materially different document from the one it hands a person, on 1.8 percent of the 436 sites in his first pass, and Time as the only site in that group serving advertising payload to agents. Zero others. His summary of it is blunt: the thing the trade press has been treating as the beginning of a wave is one publisher with two advertisers, and anyone about to build a strategy around agent-facing ad inventory has time.

This is the more useful half of the report for a site owner, and it is the half least likely to be quoted. A 1.8 percent rate across 436 sites is a real measurement of absence, and absence is what most crawler anxiety turns out to be made of. It does not make the mechanism uninteresting: a markdown twin is still a second representation free to disagree with the first, which is exactly the condition prose parity exists to detect, and the scanner detecting it still cannot tell an accident from a sale. What the figure does is put the risk in proportion. If your concern is that AI systems are being shown a doctored copy of the web, the evidence available today says that is happening on roughly one site in fifty-five, and the commercial action is somewhere else entirely.

Sample Illustrative, not a measurement of any real site.

  • 200 with the article Open The vendor reaches the content without paying. Says nothing about whether a deal exists.
  • 402 Payment Required Billed No current paid access. Could be a negotiation in progress or a lapsed renewal, not a refusal.
  • 403 or a challenge Blocked Refusal with no price offered. A commercially opposite signal to a 402.
  • Refuses a plain browser too Unmeasurable Bot defence applies to everyone, so no agent policy can be inferred. Excluded from the panel.
Four outcomes a probe can get from one origin, and what each one settles. Mechanism, not a measurement of any named site.

Why robots.txt cannot tell you who is being billed

Lane's closing observation is the one that belongs to this site, and it is not a complaint about publishers. None of this is in robots.txt. Not one of the outlets announces any of it, and the entire commercial layer, who is billed, who walks in free, which intermediary collects and what the price signal is, lives in edge configuration and response bodies that are discoverable only by impersonating a crawler and reading the status code.

That is not a failure of the file. It is the file working as specified. RFC 9309, the Robots Exclusion Protocol, published on the Standards Track in September 2022, describes itself in its abstract as a method for service owners to control how content may be accessed by automatic clients, and states plainly in its introduction that these rules are not a form of access authorization. A robots.txt file expresses a request about which paths a named crawler should fetch. It has no vocabulary for terms, prices or parties, has never had one, and the drafts that would extend the file are still drafts, as we set out in the AI preferences standard you cannot deploy yet. Asking robots.txt who is being billed is asking a permissions file a contracts question.

The vendor documentation points the same way. OpenAI's crawler documentation, read on 5 August 2026, names OAI-SearchBot, OAI-AdsBot, GPTBot and ChatGPT-User, and names exactly two mechanisms a site owner can use: robots.txt tags per bot, and its published IP ranges for verifying requests. Payment, licensing and tolls do not appear. So a publisher billing OpenAI is doing something the vendor's own integration guidance does not describe, using a status code that, as we found when reading the specification, prices crawlers with an HTTP code the spec leaves undefined.

This is the practical consequence, and it is the reason two layers decide whether AI can read your site rather than one. The file is the declared layer and the edge is the operative one, and where they disagree the edge wins every time. That divergence is measurable in both directions: it produced the finding that a robots.txt ban and a 200 response coexist on hundreds of sites, and here it produces a commercial regime that is fully deployed and completely undeclared. Neither is visible to a tool that reads only the file, which includes our own robots.txt tester when used on its own.

Where each answer is decided. The declared layer and the operative layer are different systems, and only the second one produces the status code.

What a site owner can check on their own site today

Everything above is about other people's sites, and it is reported rather than measured here. The transferable part is the method, because the same few requests answer the same question about your own origin, and almost nobody runs them.

Start with the baseline, since the probe's 17.3 percent exclusion rate is the trap. Fetch one of your own article URLs with an ordinary desktop browser user agent and record the status code. If that is already a challenge or a 401, your bot defence is refusing everyone and no per-vendor result you collect afterwards means anything. Then repeat the identical request under each named vendor token, changing nothing else, and compare status codes rather than page length. A 402, a 403 and a 200 with a truncated body are three different decisions your edge is making on your behalf, and on a managed platform or a CDN with a bot product enabled they are frequently decisions nobody at your organisation made deliberately. That is the pattern behind one in seven AI browsing agents being stopped by a bot control.

Two cautions carry over directly from Lane's corrections. Vary more than the user agent, because a single client library has a TLS fingerprint of its own and you may be measuring that instead of your policy, and do not read a refusal as a price or a price as a refusal. From outside, a scanner carrying no payment header cannot distinguish a page that costs money from a page that is closed, which is a limit we have stated about our own tooling and not a solvable one.

For the file half of the question, the robots.txt tester resolves your file per crawler token, which answers what you have declared. For the response half, what GPTBot sees shows what actually comes back. The two answering differently is the finding, not an error, and it is worth checking after any edge change because a robots.txt edit does not take effect when you save it either. If you are trying to be cited rather than trying to be paid, the platform guide for getting cited in ChatGPT starts from the same place: a client that is refused at the door cites nothing, and no amount of generative engine optimization further up the stack changes that.

  • Human baseline first Fetch as a desktop browser user agent. A challenge here means nothing else you measure is interpretable.
  • One request per vendor token Change only the user agent, and record the status code rather than the body length.
  • Two independent clients A single library's TLS fingerprint can produce a result that has nothing to do with your policy.
  • Separate 402 from 403 A bill and a block are opposite commercial signals. Collapsing them inverts the finding.
  • Compare file against response What robots.txt declares and what the edge returns are different answers to different questions.
  • Price against block, from outside Not resolvable by any unauthenticated scanner. A client with no wallet gets the same refusal either way.
The checks that reproduce this method on one origin. Every one of them is a request you can send yourself.

Related

Common questions

Does a 402 response mean a publisher has refused an AI company?

No. Lane's report is explicit that a 402 means no current paid access rather than refusal, so a billed vendor may be in live negotiation or on a renewal that lapsed. A 403 or a challenge with no price attached is the refusal signal, and the report counts those separately: 30.7 percent of the 153 clean-baseline outlets block outright while 17.0 percent gate at least one vendor behind a price.

Can I find out which AI companies a publisher has a deal with by reading its robots.txt?

No. The report states that none of the publishers in the panel announce any of this in robots.txt. Under RFC 9309 a robots.txt file expresses which paths a named crawler is asked to fetch, and the specification says its rules are not a form of access authorization, so the file carries no vocabulary for parties, terms or prices. The commercial decision is made in edge configuration and is visible only in the response.

Has Lantad measured how many sites charge AI crawlers for access?

No. Lantad operates no publisher panel and publishes no figure for toll deployment. Every number in this post is reported from Kin Lane's probe published at API Evangelist on 4 August 2026. Lantad measures what a named crawler receives when it requests a URL and how that compares with what a browser renders, and from outside it cannot separate a priced page from a blocked one, because a scanner sends no payment header.

Are many publishers now serving ads to AI agents?

Not on the evidence available. The probe found differential serving on 1.8 percent of 436 sites, and Time was the only publisher in the panel serving advertising payload to agents. The commercial activity the probe did find at scale was tolling rather than advertising, with fourteen gated outlets collecting through TollBit and twelve running a first-party toll.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.