BlogFindings

Meta took the crawl volume and ChatGPT kept 80 to 88 percent of the referrals

DataDome's press release of 16 July 2026 states that its network processed 17.7 billion AI agent requests between April and June, up from 12.2 billion in the previous quarter, with Meta's two agents taking the majority of that traffic. In the same release, ChatGPT-User's absolute request volume fell 6 percent while ChatGPT held 80 to 88 percent of all AI referrals every month. Crawl volume and referral value are not the same population, and the tokens are not interchangeable.

17 min read Lantad

DataDome, a bot and agent trust management vendor, published a quarterly analysis on 16 July 2026 that puts numbers on the gap. Its press release, headed The AI Traffic Report Q2 2026: Agentic Traffic Surged 45%, With Meta Taking the Lead, describes AI traffic across its network from April to June 2026. Two of its findings point in opposite directions: the agents driving request volume grew sharply and belong to Meta, while the product driving referral traffic is ChatGPT, whose own user agent fell in absolute volume over the same quarter. Lantad has measured none of this and cannot: the figures come from one vendor's network, which is not a place we can observe. What follows reports the release, then covers the part that is checkable on your own site, which is whether the tokens in your robots.txt are the ones that matter for each of those three questions. The report page itself at datadome.co/threat-research/ai-traffic-report-q2-2026/ returned HTTP 403 to two requests from this machine on 13 August 2026, one carrying a browser user agent and one carrying ours, so the figures below are read from the press release as distributed by Business Wire at www.businesswire.com/news/home/20260716617247/en/ rather than from the report. Where the release states a percentage and no absolute count, this post says so rather than filling the gap.

In short

  • DataDome's press release of 16 July 2026 states that its network processed 17.7 billion AI agent requests in the second quarter of 2026, up from 12.2 billion in the first, which it describes as a 45 percent increase quarter on quarter.
  • The same release states that Meta-ExternalAgent grew 74 percent and Meta-WebIndexer grew 163 percent quarter on quarter, and that Meta's two agents together now account for the majority of all AI agent traffic on that network, a position they did not hold in the first quarter.
  • It also states that ChatGPT-User, the top agent in the first quarter, declined 6 percent in absolute request volume, while ChatGPT commanded 80 to 88 percent of all AI referrals every month and grew referrals 17 percent quarter on quarter.
  • OpenAI's crawler documentation, read on 13 August 2026, states that ChatGPT-User is not used for crawling the web in an automatic fashion, and that because these actions are initiated by a user, robots.txt rules may not apply.
  • Lantad measured none of this. Every figure here is DataDome's, taken from its press release of 16 July 2026, and it describes traffic across the customer base of one bot defence vendor rather than the web.
What the release statesFigureBasis
AI agent requests on DataDome's network, second quarter17.7 billionApril to June 2026
AI agent requests, first quarter12.2 billionThe stated comparison base
AI agent traffic, quarter on quarterup 45 percentConsistent with the two totals
Meta-ExternalAgent, quarter on quarterup 74 percentPercentage only, no count published
Meta-WebIndexer, quarter on quarterup 163 percentPercentage only, no count published
ChatGPT-User, absolute request volumedown 6 percentDescribed as the top agent in the first quarter
ChatGPT share of all AI referrals80 to 88 percentMonthly, referrals rather than crawl requests
ChatGPT referrals, quarter on quarterup 17 percentReferrals rather than crawl requests
MCP traffic, peak dayapproaching 500,000 requestsDescribed as a measurable network signal
DataDome customers with agent trust policies54 percentCustomers, not sites on the open web
Every figure DataDome's press release of 16 July 2026 states about the quarter, read on 13 August 2026. Reported from DataDome, not measured by Lantad. Where the release gives a percentage and no count, the count is not published.

What DataDome reported for the second quarter of 2026

The headline number is network wide. DataDome's release states that its network processed 17.7 billion AI agent requests in the second quarter of 2026, up from 12.2 billion in the first, and describes that as AI agent traffic surging 45 percent quarter on quarter. The arithmetic is consistent: 17.7 divided by 12.2 is 1.45, so the percentage and the two totals agree, which is worth checking on any release that gives you both and is not always true.

The composition is the part that changed. The release states that Meta-ExternalAgent grew 74 percent quarter on quarter and Meta-WebIndexer grew 163 percent, and that together Meta's two agents now account for the majority of all AI agent traffic on the network, a dominance that did not exist in the first quarter. Both names will be familiar to anyone who read what Meta's own crawler documentation names: Meta-WebIndexer is the one Meta describes as serving Meta AI search, which makes it the more consequential of the two for anybody thinking about AI visibility rather than about bandwidth.

Then the finding this post is named after. The release states that ChatGPT-User, the top agent in the first quarter, declined 6 percent quarter on quarter in absolute request volume, and in the same breath that ChatGPT remains the overwhelming leader in referral traffic, commanding 80 to 88 percent of all AI referrals every month and growing 17 percent quarter on quarter. DataDome's own framing of this is that crawl volume and referral value are moving in opposite directions, and its VP of Threat Research, Jérôme Segura, is quoted saying that ChatGPT is driving more referral value with fewer crawls.

Two smaller figures round it out. The release states that MCP traffic is now a measurable network signal, with peaks approaching 500,000 requests per day and clear daily usage cycles. That is the only line in the release about a protocol rather than a crawler. It also states that 54 percent of DataDome customers have already implemented agent trust policies. Note the denominator on that last one: it is DataDome's customers, who are by definition sites that bought bot defence, and it says nothing about the web.

What the release does not contain is as important as what it does. There is no absolute request count for any individual agent, no statement of how many domains the network covers, no definition of what counts as an AI agent request, and no description of how a referral is attributed. Those are the four things you would need to reproduce any of it. This post is not going to invent them.

  • Network total Stated 17.7 billion AI agent requests in the second quarter, against 12.2 billion in the first.
  • Per agent growth Stated as percentages Meta-ExternalAgent up 74 percent, Meta-WebIndexer up 163 percent, ChatGPT-User down 6 percent. No absolute counts.
  • Referral share Stated ChatGPT at 80 to 88 percent of all AI referrals every month, growing 17 percent quarter on quarter.
  • Domain coverage Not stated The release names no number of sites or domains behind the network total.
  • Classification method Not stated How a request is judged to be an AI agent request is not described in the release.
  • Referral attribution Not stated How a visit is credited to an AI product, and over what window, is not described.
What DataDome's press release of 16 July 2026 states, and what it does not state, read on 13 August 2026. The second column is the reason the figures above carry the qualifiers they do.

Why crawl volume and referral traffic are different populations

The temptation with a number like 17.7 billion is to treat it as a measure of AI interest in the web, and then to treat a fall in one agent's volume as a fall in that agent's importance. The second quarter figures show why that does not follow, and the reason is in the vendor documentation rather than in the traffic.

OpenAI's crawler documentation, read on 13 August 2026, names four agents and gives each a different job. It states that OAI-SearchBot is for search, that GPTBot is used to make its generative AI foundation models more useful and safe, that OAI-AdsBot validates the safety of web pages submitted as ads, and that OpenAI also uses ChatGPT-User for certain user actions in ChatGPT and Custom GPTs. Two sentences on that page decide how the 6 percent should be read. The first is that ChatGPT-User is not used for crawling the web in an automatic fashion. The second is that because these actions are initiated by a user, robots.txt rules may not apply.

Read those together and ChatGPT-User is not a crawler in the sense the word is normally used. It is a fetch that happens because a person asked something, so its volume tracks how often ChatGPT decides a live page is needed to answer, which is a product decision that can change in a release note. A 6 percent fall in that number is a statement about retrieval behaviour, not about how many people used the product or how many of them clicked a link. The referral figure in the same release moved the other way, and both can be true at once without contradiction because they count different events.

The population confusion runs in the other direction too. A crawler request is not a reader. Between a request arriving and a person landing on your site sit several independent steps, each of which can fail on its own: the response has to carry your text rather than an empty shell, the content has to enter an answer system, the answer has to cite you, and somebody has to click. We have written before that the page that earns a citation is often not the page that receives the visit, and separately that most AI citations point at other companies' domains rather than at the brand being asked about. Neither of those failures is visible in a request count, and no amount of crawl volume compensates for either.

This is also why generative engine optimisation advice that begins and ends with a robots.txt edit tends to disappoint. The file governs the first step. The other steps are governed by what your server returns, what your page contains, and what the model does with it, and those are separate systems that fail separately.

Sample Illustrative, not a measurement of any real site.

The steps between one AI agent request and one person arriving on your site. Each is a separate system with its own failure mode, and a request count reports only the first. Illustrative of the mechanism, not a measurement of any real site.

The tokens driving volume are not the tokens driving citations

Every vendor that publishes more than one crawler name publishes them with different jobs, and the jobs map onto the three questions at the top of this post. Read on 13 August 2026, three vendor pages say so in their own words.

Anthropic's documentation, which carries a date of 7 April 2026, names three. It states that ClaudeBot collects web content that could potentially contribute to training, that Claude-User supports Claude AI users and may access websites when individuals ask questions, and that Claude-SearchBot navigates the web to improve search result quality by analysing online content to enhance the relevance and accuracy of search responses. Perplexity's page names two and is explicit about the boundary: it describes PerplexityBot as designed to surface and link websites in search results and states it is not used to crawl content for AI foundation models, and describes Perplexity-User as supporting user actions and likewise not used for training. Neither the Perplexity page nor OpenAI's carries a last updated date, which is worth noting when you are deciding how current your notes are.

The practical consequence is that a robots.txt group written against the wrong name achieves the opposite of what was intended. Disallowing the search token and allowing the training token is a coherent set of rules that nobody would choose deliberately, and it is easy to arrive at by copying a list. It is also, on the numbers in the DataDome release, the more expensive mistake in one direction than the other: the tokens that generated the volume growth are Meta's, and the product that generated the referrals is ChatGPT, so a rule tuned to reduce bandwidth from the first has no bearing on the second.

This assumes you have two names to write rules against, and frequently you do not. We counted this in the registry and found that six of nine vendors publish exactly one crawler token, which means the standard advice to allow search and block training cannot be expressed at all for most of them. Where a vendor publishes a robots.txt product token with no matching user agent string, the problem is worse again, because those tokens never appear in your logs and a log search for them returns a zero that means nothing.

One more rule decides what happens to a name you have not written a group for. RFC 9309 specifies that a crawler matching no group of its own falls to the group headed by an asterisk, so every token you have never heard of is already governed by whatever that group says. Any agent named after your file was last edited is in exactly that position. Whatever your wildcard group currently says is your published answer for every agent that arrives next quarter.

VendorTokenWhat the vendor's page says it is for
OpenAIOAI-SearchBotFor search
OpenAIGPTBotTo make generative AI foundation models more useful and safe
OpenAIChatGPT-UserCertain user actions in ChatGPT and Custom GPTs, not automatic crawling
OpenAIOAI-AdsBotValidating the safety of pages submitted as ads on ChatGPT
AnthropicClaudeBotCollecting web content that could contribute to training
AnthropicClaude-UserSupporting Claude AI users when individuals ask questions
AnthropicClaude-SearchBotImproving search result relevance and accuracy
PerplexityPerplexityBotSurfacing and linking websites in search results, not training
PerplexityPerplexity-UserSupporting user actions, not crawling or training
What three vendors say each of their tokens is for, quoted in summary from their own documentation pages read on 13 August 2026. Reported from vendor documentation, not measured by Lantad.

What a network wide share does not tell you about your site

A figure like 80 to 88 percent of all AI referrals invites a substitution: treat it as the share you should expect. It is not, and three separate things stand between the two.

The first is the sample. DataDome's network is its customer base, which is a population of sites that decided bot traffic was a problem worth paying to manage. That is a selected group in exactly the direction that matters, and the release itself gives the clue, since 54 percent of those customers already run agent trust policies. Whatever the mix of AI traffic looks like across sites that have deployed bot defence and agent rules, it is not a random sample of the web, and the release makes no claim that it is.

The second is what a per agent count actually counts. Every per crawler number anywhere, including ours, counts requests that said they were that crawler, and we have written about why a user agent is a claim rather than an identity. Verifying the claim needs a reverse DNS or published IP range check. A bot defence vendor is better placed than most to do that verification, and the release does not say whether these counts are verified or self declared, so the honest reading is that the direction of each movement is more solid than any individual figure.

The third is that a share of referrals is a ratio, not a volume. ChatGPT holding 80 to 88 percent of AI referrals is compatible with AI referrals being a large source of traffic or a rounding error for any given site, because the denominator is all AI referrals rather than all visits. The release does not state what proportion of total traffic AI referrals represent, and anyone quoting the 80 to 88 percent figure as evidence about traffic volume is adding something the source does not say.

None of this makes the report less useful. Direction of travel across 17.7 billion requests is real information and there is no public dataset that would let an individual site reconstruct it. It means the number belongs in the sentence it was published in. Our own methodology page sets out the same limits for what we publish, and the research page states which of our figures rest on a sample large enough to quote.

  • Selected sample Customers, not the web The network is one vendor's customer base, and 54 percent of it already runs agent trust policies.
  • Self declared agents Verification not stated Per agent counts count requests carrying that name. The release does not say whether they were verified.
  • Share, not volume Ratio only 80 to 88 percent is a share of AI referrals, and the release states no proportion of total traffic.
  • Direction of travel Usable Growth and decline across 17.7 billion requests is real information about the quarter.
Three gaps between a network wide figure and your own site. Each one is a reason to quote the figure with its qualifier rather than as a benchmark.

What to check on your own site this week

None of the above changes what you should do, but it changes the order. Four checks, each of which produces an answer about your site rather than about a network.

Start with the names your robots.txt actually mentions, because that file is the only part of this you author. Meta-WebIndexer, the fastest growing of the agents the release names, is a recent addition to Meta's documented tokens, so a file written before it was published cannot name it and the wildcard group is answering for it. Read your file and decide deliberately whether that is the answer you want to give it, then do the same for the search tokens and the user triggered tokens of each vendor above. Our robots.txt tester evaluates a file per crawler so you can see which group each name lands in, and the AI crawler reference lists the tokens themselves.

Second, confirm the file is the layer that answers. It usually is not on its own. What decides whether a crawler is served is the response your infrastructure returns, and those two layers can disagree without anything reporting the difference. The strongest evidence that this is common rather than theoretical is a measurement we covered earlier: 234 of 592 sites that ban GPTBot in robots.txt served it a 200 anyway. A file that says one thing while the origin does another is not a policy, it is two policies.

Third, separate the three questions in your own reporting. Count crawl requests by verified agent, count whether those requests received readable text, and count referrals from assistant hosts, and keep them in different columns. A single AI traffic number that adds a crawler hit to a human visit produces a quantity with no meaning, and the temptation to build it is strongest when the crawler number is large.

Fourth, be careful about what blocking buys you. A canary study we covered found that a robots.txt block did not stop 12 of 18 AI chatbots from returning site content, largely because a search engine's crawler had already supplied it. If your goal is bandwidth, blocking the high volume agents is a reasonable lever and the release tells you which they are. If your goal is to keep content out of answers, the evidence that robots.txt achieves it is weak. And if the goal is citations, the file to check is the one governing the search and user triggered tokens, which is the same set covered in our guide to getting cited in ChatGPT.

The last observation is about the MCP figure. Peaks approaching 500,000 requests a day is small next to 17.7 billion, and it is the only line in the release describing a channel that did not exist as a measurable signal a year ago. We have written about what it takes for a site's tools to be discoverable to an agent, and the short version is that discovery is the unsolved part. A traffic report is a lagging indicator of where that goes next, which is the right way to read all of these numbers: a description of the quarter that ended, not a forecast of the one you are in.

  • Every vendor token named explicitly in robots.txt Meta-WebIndexer and the user triggered tokens are absent from most files, so the wildcard group answers for them.
  • Robots.txt and the response agree The file states intent, the origin and CDN decide the outcome, and nothing reports a disagreement.
  • Crawl hits, readable responses and referrals counted separately Three questions, three columns. A single AI traffic total adds a machine to a person.
  • Blocking chosen for a stated goal Bandwidth, exclusion from answers and citations are three different goals, and robots.txt serves them unequally.
Four checks that produce an answer about your own site. None of them requires the figures in this post, and none of them is settled by a request count.

Written by

Lantad

Published .

Three different questions get answered with the same number. How much does an AI system fetch from your site, can it read what it fetched, and does any of it send a person back to you. Server logs answer the first cheaply, which is why it is the one everybody quotes, and it is the one with the weakest relationship to the outcome most site owners actually want.

Common questions

How many AI agent requests did DataDome report for the second quarter of 2026?

17.7 billion. DataDome's press release of 16 July 2026 states that its network processed 17.7 billion AI agent requests in the second quarter of 2026, up from 12.2 billion in the first quarter, which it describes as a 45 percent increase. The release covers April to June 2026 and does not state how many domains sit behind that total.

Why did ChatGPT-User request volume fall while ChatGPT referrals grew?

They count different events. OpenAI's crawler documentation, read on 13 August 2026, states that ChatGPT-User is not used for crawling the web in an automatic fashion and is used for certain user actions in ChatGPT and Custom GPTs, so its volume reflects how often the product fetches a live page to answer. Referrals count people arriving on a site. DataDome's release states the first fell 6 percent quarter on quarter and the second grew 17 percent, and both can be true at once.

Should I block Meta-WebIndexer in robots.txt?

It depends which goal you are serving, and it is a decision rather than a default. Meta describes Meta-WebIndexer as serving Meta AI search, so blocking it is a decision about appearing in that surface rather than only about bandwidth. DataDome's release states it grew 163 percent quarter on quarter on that network. If the name is not in your file today, RFC 9309 means your wildcard group is already answering for it.

Did Lantad measure any of these figures?

No. Every traffic figure in this post is DataDome's, taken from its press release of 16 July 2026, and describes traffic on that vendor's network rather than the web. The vendor documentation quotes are read from OpenAI, Anthropic and Perplexity pages on 13 August 2026. Lantad measures whether AI crawlers can read a given site, which is a different question from how often they arrive.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.