BlogFindings
Meta AI crawler: 6 of 1,065 robots.txt files named the token Meta ties to citation
Lantad requested the robots.txt of all 1,419 hostnames in this repository's corpus seeds on 24 September 2026 as LantadBot, and 1,065 returned a plain-text file. 117 named Meta-ExternalAgent, the training crawler. 6 named Meta-WebIndexer, the crawler Meta's own documentation says has to be allowed for Meta AI to cite and link to a site. All six of those refused it.
This one does. On 24 September 2026 Lantad requested https://host/robots.txt for all 1,419 hostnames in this repository's corpus seeds as LantadBot, following redirects, and 1,065 of them returned a plain-text file rather than an error, an HTML page or nothing at all. Every count below is over those 1,065 files, parsed and evaluated with the same group-selection code the robots.txt tester runs, so a verdict here is the verdict the product would give.
Six files named Meta-WebIndexer. One hundred and seventeen named Meta-ExternalAgent. The gap between those two numbers is the subject, and it turns out to be the widest version of a gap that exists for every vendor that splits training from search.
In short
- Lantad read 1,065 plain-text robots.txt files on 24 September 2026. 117 named Meta-ExternalAgent and 6 named Meta-WebIndexer, the two tokens Meta's documentation separates into training and Meta AI search.
- All six files naming Meta-WebIndexer on 24 September 2026 disallowed it at the site root: congress.gov, eltiempo.com, news24.com, theglobeandmail.com, amazon.com and qcitymetro.com. Not one file named it in order to allow it.
- Refusing the Meta AI crawler that trains did not refuse the one that searches. 58 of the 1,065 files blocked Meta-ExternalAgent at the root and 42 of those 58 still allowed Meta-WebIndexer there, 40 of them on a wildcard group that never names Meta.
- A naming count is not a preference count. 42 of the 117 files naming Meta-ExternalAgent named it in a group shared with the wildcard token, all 42 allowed it at the root, and 35 of those 42 were Wix or Squarespace sites shipping a platform default.
- The awareness gap between a vendor's training token and its search token was widest for Meta in this sample: 171 files named GPTBot against 79 for OAI-SearchBot, 157 named ClaudeBot against 38 for Claude-SearchBot, and 117 named Meta-ExternalAgent against 6 for Meta-WebIndexer.
| Token | What Meta's page says it is for | Files naming it | Share of 1,065 |
|---|---|---|---|
| Meta-ExternalAgent | Training foundation AI models, and indexing content directly | 117 | 11.0% |
| Meta-ExternalFetcher | Fetching links a user asked for; may bypass robots.txt | 27 | 2.5% |
| FacebookExternalHit | Link previews in Facebook, Instagram and Messenger | 25 | 2.3% |
| Meta-WebIndexer | Meta AI search; allowing it helps Meta AI cite and link to you | 6 | 0.6% |
| Meta-ExternalAds | Improving advertising and other business products | 1 | 0.1% |
Which Meta AI crawler tokens do robots.txt files actually name?
134 of the 1,065 files name at least one of Meta's five. Meta-ExternalAgent accounts for almost all of that: 117 files. Meta-ExternalFetcher appears in 27, FacebookExternalHit in 25, Meta-WebIndexer in 6 and Meta-ExternalAds in 1, which was slate.com. The five tokens are not equally known, and the ordering is not the ordering Meta's own descriptions would suggest if a reader were choosing which ones to care about.
Set that against the wider AI crawler picture in the same files and the shape gets clearer. 202 of the 1,065 files name at least one of the fourteen non-Meta AI crawler tokens this scanner tracks, which is the population that has thought about AI crawlers at all rather than the population that has thought about Meta. 82 of those 202 name no Meta token whatsoever. So among sites that have demonstrably read something about AI crawling and acted on it, two in five wrote nothing about Meta, and a much larger fraction wrote about the wrong Meta token.
The sample matters for how far any of this generalises, and it is worth being exact rather than reassuring. These are the 1,419 hostnames this repository seeds its corpus from, grouped into industry categories and website-builder platforms, and they were chosen to span the web rather than to be drawn at random from it. Government, education, healthcare, news, SaaS, ecommerce, travel and finance sites sit alongside Framer, Webflow, Wix, Squarespace, Shopify and WordPress builds. How that frame is constructed and what it will not support is written up in the crawlability study. Read the numbers below as a measurement of 1,065 named sites on one date, not as a percentage of the web.
One more scoping note before the counts. A token appearing in a file tells you somebody typed it. It does not tell you what happens when the crawler arrives, because that depends on which group the token sits in and what rules hang off that group. The two questions come apart badly here, and the section on naming against ruling is where they are separated. The AI crawlers reference lists tokens rather than companies for exactly this reason: the unit a robots.txt file operates on is the token, and a vendor is not a unit at all.
Six files named Meta-WebIndexer, and all six said no
The six are congress.gov, eltiempo.com, news24.com, theglobeandmail.com, amazon.com and qcitymetro.com. Four of the six are news organisations, one is the United States legislative information site and one is a retailer. The news slice of the sample leans the same way on the better known token: of the 64 news files, 30 named Meta-ExternalAgent and 29 of those 30 blocked it at the root. Publishers read crawler documentation. Most other people do not.
Every one of the six wrote Disallow: / under the token. Not one file in 1,065 named Meta-WebIndexer in order to allow it. That is a finding about what the files say and not a claim about anybody's intent, and it is worth holding at that width: six site owners, or six vendors acting for them, learned the name of Meta's search crawler and used it to refuse. If the sentence on Meta's page is accurate, those six have told Meta AI not to cite them, and four of them publish news for a living.
The other 1,049 allow it, and almost none of them chose to. 16 files block Meta-WebIndexer at the root: the six above, plus ten where a wildcard group carrying Disallow: / swept it up along with everything else. The remaining 1,049 allow it by default, which is what happens in the absence of a matching rule and what RFC 9309 specifies when no group matches at all. Silence is permission in this format, so most of those 1,049 sites are in the state Meta's documentation recommends without having decided to be.
Whether being allowed produces a citation is a different question and this measurement does not touch it. Access is necessary and not sufficient, which is the distinction the GEO category collapses more often than any other, and it is why seeing what a crawler receives and watching what the models actually say are two purchases rather than one. What these files support is narrower and still worth knowing: on 24 September 2026 the access condition Meta names in its own documentation was met on 1,049 of 1,065 sites and refused on 16, and six of those 16 refused it deliberately.
| Host | Rule under the token | Verdict at / |
|---|---|---|
| congress.gov | Disallow: / | Blocked |
| eltiempo.com | Disallow: / | Blocked |
| news24.com | Disallow: / | Blocked |
| theglobeandmail.com | Disallow: / | Blocked |
| amazon.com | Disallow: / | Blocked |
| qcitymetro.com | Disallow: / | Blocked |
Blocking Meta's training crawler did not block its search crawler
58 of the 1,065 files disallow Meta-ExternalAgent at the site root. That is a preference expressed in the mechanism Meta asks for, and it is the decision most coverage of Meta and AI crawling is about. 42 of those 58 files allow Meta-WebIndexer at the same path.
The cause is group selection, and it catches people who have read enough about robots.txt to be confident. Rules do not accumulate across groups: a crawler picks the single most specific group whose token matches it, and only that group's rules apply. A file with a Meta-ExternalAgent group carrying Disallow: / has written a rule that Meta-ExternalAgent obeys and Meta-WebIndexer never reads, because Meta-WebIndexer is not in that group. It falls through, and on 40 of those 42 files the wildcard group decided the verdict. On the other two, no group matched and the default applied. This is the mechanic behind 559 of 581 lost pages being closed by a rule that never named the crawler, running in the opposite direction: there it blocked more than the author meant, here it blocks less.
Sites in that 42 include fema.gov, unesco.org, coursera.org, nature.com, drugs.com and jamanetwork.com. Each wrote a rule refusing Meta's training crawler and left Meta's search crawler with root access. Whether that is what they wanted is not something a robots.txt file records, and it may well be: refusing training while permitting search is a coherent position and an increasingly common one among publishers who want the citation and not the corpus. The point is that none of those 42 files says so. The permission is a consequence of specificity rather than a sentence anybody wrote, and the same asymmetry between training rules and search rules appeared when we counted Google-Extended against Googlebot.
Checking your own file is easy and eyeballing it is not, which is the practical takeaway. Evaluate the file once per token rather than once per vendor, and read the group that matched rather than the rule you remember writing. Getting the ordering of those two steps wrong is a recurring failure that AI crawler detection is an ordering problem covers at length, and it produces confident wrong answers rather than obvious ones.
Sample Illustrative, not a measurement of any real site.
Flow: One robots.txt to Group: Meta-ExternalAgent; One robots.txt to Group: wildcard; Group: Meta-ExternalAgent (Disallow /) to Meta-ExternalAgent; Group: wildcard (no rule, allowed) to Meta-WebIndexer.
Naming a crawler is not the same as ruling on it
117 files name Meta-ExternalAgent, and that number overstates how many people decided anything. 42 of the 117 name it inside a group that also carries the wildcard token, which is the shape a platform default takes: a stack of twenty or thirty User-agent lines closed by an asterisk, with one set of rules underneath aimed at nobody in particular. All 42 allow Meta-ExternalAgent at the root, because the rules attached to that group are the ordinary housekeeping disallows a hosted site ships with.
35 of the 42 are Wix or Squarespace sites. One of them lists CCBot, ClaudeBot, cohere-ai, DuckAssistBot, FacebookBot, Google-Extended, GPTBot, img2dataset, Meta-ExternalAgent, omgili, Quora-Bot, TikTokSpider, YouBot and a dozen more as consecutive User-agent lines, closes the stack with an asterisk, and then disallows /config, /search and a few account paths. Every AI crawler in that list is allowed everywhere that matters. The file reads like a considered AI policy and functions as a default, which is the mirror image of what we found when 79 of 86 Drupal files named no AI crawler at all and 68 of those turned out to be Drupal's own shipped file, unedited.
The 75 files that name Meta-ExternalAgent in a group of its own are the real signal, and 51 of those 75 block it at the root. That is the population worth reporting as a preference: roughly 5 percent of the sample refused Meta's training crawler in a rule written for Meta. The distance between 117 and 51 is not anybody being dishonest. It is that a count of names is a count of strings, which is why 2,209 tokens across 1,004 files told us considerably less than evaluating those files did.
The correction applies to every naming count in this post and to most published elsewhere. GPTBot is named in 171 files and shares a wildcard group in 43 of them. ClaudeBot is named in 157 and shares one in 43. Google-Extended is named in 145 and shares one in 44. Between a quarter and a third of every token's naming count is a platform default expressing no preference at all, so a headline of the form "N percent of sites now block X" that does not separate the two is partly reporting the shipping defaults of three website builders. The same caution applies to our own earlier count of 16 of 140 files naming OAI-SearchBot.
| Token | Files naming it | Sharing the wildcard group | Blocked at the root |
|---|---|---|---|
| GPTBot | 171 | 43 | 72 |
| ClaudeBot | 157 | 43 | 61 |
| Google-Extended | 145 | 44 | 53 |
| Meta-ExternalAgent | 117 | 42 | 51 |
| Meta-ExternalFetcher | 27 | 0 | 19 |
| Meta-WebIndexer | 6 | 0 | 6 |
Meta's training and search gap is the widest of the three
Three vendors in this sample publish both a training token and a separate search token, so the comparison runs three times on the same files. OpenAI documents GPTBot for the generative models and OAI-SearchBot for surfacing sites in ChatGPT search on its crawler page. Anthropic documents ClaudeBot alongside Claude-SearchBot and Claude-User in its crawler support article. Meta documents Meta-ExternalAgent and Meta-WebIndexer.
In the 1,065 files, GPTBot is named 171 times against 79 for OAI-SearchBot, a ratio of about 2.2 to 1. ClaudeBot is named 157 times against 38 for Claude-SearchBot, about 4.1 to 1. Meta-ExternalAgent is named 117 times against 6 for Meta-WebIndexer, about 19.5 to 1. The gap between a vendor's training token and its search token exists everywhere, and it is roughly an order of magnitude wider for Meta than for OpenAI.
The likely explanation is dull and these files cannot prove it: a token documented and written about for longer accumulates more mentions, and the robots.txt blocklists people copy from each other are years old in places. In an earlier Lantad count of 1,004 files, 120 still named a token Anthropic no longer documents, which is what a copied list looks like after the vendor has moved on. A file is a record of when somebody last cared rather than of what a vendor currently publishes, and that is true of every number in this post.
What follows for a reader is the same whatever the cause. If you wrote AI crawler rules at any point, the list you wrote them from is probably older than the documentation, and the tokens most likely to be missing are the search-side ones, which are also the ones that decide whether you can be cited. Open the vendor pages rather than a roundup: OpenAI's and Anthropic's are linked above, Meta's is at developers.facebook.com/documentation/sharing/webmasters/web-crawlers, and the guides for getting cited in ChatGPT, in Claude and in Perplexity set out what each vendor honours.
Training tokens named
- GPTBot: 171 files
- ClaudeBot: 157 files
- Meta-ExternalAgent: 117 files
Search tokens named
- OAI-SearchBot: 79 files, 2.2 to 1
- Claude-SearchBot: 38 files, 4.1 to 1
- Meta-WebIndexer: 6 files, 19.5 to 1
What this scanner still does not check, seven weeks later
The 8 August post recorded that Lantad's crawler registry carried one of Meta's five tokens. That array is BOT_REGISTRY in core/src/bots.ts, it still holds 15 tokens on 24 September 2026, and exactly one of them is still Meta's: Meta-ExternalAgent, with its purpose recorded as training. Meta-WebIndexer, Meta-ExternalFetcher, Meta-ExternalAds and FacebookExternalHit are still absent. Seven weeks after publishing the gap, a scan run today reports whether Meta's training crawler can reach a page and says nothing about the crawler Meta ties to citation.
Recording that twice is less comfortable than recording it once, and it is the honest state of the product on the day this measurement makes the gap most visible. It is also a configuration decision rather than a measurement, which is the distinction that matters when reading any registry count: the list is something somebody maintains, its contents are a choice about what to check, and a site passing 15 of 15 has passed the checks we run rather than every check that exists. The general form of that caveat is in the methodology, and the specific one is in the post that first reported it.
The 27 files naming Meta-ExternalFetcher are the other half of the problem, and this half is Meta's. 19 of those 27 disallow it at the root. Meta's page states that this crawler may bypass robots.txt because its fetches are requested by a user, so 19 site owners have written a rule against a client whose vendor said in advance that the rule is advisory. That is fairly read as a preference recorded for the record rather than an enforcement mechanism, and it is how user-triggered and agent-triggered fetches increasingly work across the industry. It is still worth knowing before counting that rule as protection.
None of this argues for writing more rules. It argues for knowing which of the rules you already have are doing something. Evaluate the file once per token, read the group that actually matched, and take the vendor documentation rather than a copied list as the source of which tokens exist. And check that the file can be fetched at all before trusting any of it, because 92 of 1,056 sites refused GPTBot the file itself, and a rule nobody can read is not a rule.
- Meta-ExternalAgent In the registry, purpose recorded as training. One of the 15 tokens evaluated.
- Meta-WebIndexer Not checked. The token Meta's page ties to citation in Meta AI responses. Reported missing on 8 August 2026 and still missing.
- Meta-ExternalFetcher Not checked. Named by 27 files in this sample, 19 of which disallow it at the root.
- Meta-ExternalAds Not checked. Named by one file in this sample, slate.com.
- FacebookExternalHit Not checked. Named by 25 files, and it predates the AI crawler question entirely.
Lantad
Published .
On 8 August 2026 this blog published a correction: Meta documents five crawler tokens, not the one we had reported a week earlier, and the most consequential of the five is Meta-WebIndexer. Meta's webmaster page describes it as navigating the web to improve Meta AI search result quality, and then says the part that turns a token into a decision: allowing Meta-WebIndexer in your robots.txt helps Meta cite and link to your content in Meta AI's responses. That post was about what the vendor publishes. It did not ask what site owners had written down.
Common questions
What is the Meta AI crawler's user agent?
There are five, not one. Meta's webmaster documentation names FacebookExternalHit, Meta-WebIndexer, Meta-ExternalAds, Meta-ExternalAgent and Meta-ExternalFetcher, each with a UA string of the form token/1.1. Meta-ExternalAgent is the training crawler and the one most robots.txt files name: 117 of the 1,065 files Lantad read on 24 September 2026. Meta-WebIndexer is the Meta AI search crawler, named by 6 of those files, and Meta's page states that allowing it helps Meta cite and link to your content.
How do I block the Meta AI crawler in robots.txt?
Write one group per token you want to refuse, because robots.txt applies the single most specific matching group rather than every group that could apply. A group headed User-agent: meta-externalagent with Disallow: / refuses training and leaves Meta-WebIndexer untouched, which is exactly what 42 of the 58 files blocking Meta-ExternalAgent in Lantad's 24 September 2026 sample did. Meta's page also states that Meta-ExternalFetcher may bypass robots.txt because its fetches are user-triggered, and that a robots.txt change can take up to 24 hours to take effect.
Does blocking Meta-ExternalAgent stop Meta AI from citing my site?
Not on its own, by Meta's own account. Meta's documentation ties citation in Meta AI responses to allowing Meta-WebIndexer, a different token in a different group. Of the 1,065 plain-text robots.txt files Lantad read on 24 September 2026, 58 blocked Meta-ExternalAgent at the root and 42 of those still allowed Meta-WebIndexer there. Access is a necessary condition rather than a sufficient one, so an allowed crawler is not a promise of a citation and this measurement does not claim it is.
How many sites block the Meta AI crawler?
In Lantad's sample of 1,065 plain-text robots.txt files read on 24 September 2026, 58 disallowed Meta-ExternalAgent at the site root and 16 disallowed Meta-WebIndexer there. 117 files named Meta-ExternalAgent, but 42 of those named it inside a group shared with the wildcard token, which is the shape of a website builder's default file and blocks nothing. The sample is this repository's corpus seeds rather than a random sample of the web, so treat it as a measurement of 1,065 named sites on one date.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.