BlogFindings
DuckAssistBot: 63 of 1,063 robots.txt files named it, and 48 of them let it in
DuckDuckGo documents that DuckAssistBot fetches pages in real time for answers that cite their sources, that the data never trains a model, and that disallowing it costs a site nothing in search. That makes it the cheapest citation crawler on the web to refuse, and the only lever is the product token in robots.txt. Lantad requested the robots.txt of all 1,419 hostnames in this repository's two committed corpus seed files on 7 October 2026. 1,063 returned a file that parsed as plain text, 63 of those named DuckAssistBot, and 48 of the 63 left the root open to it.
Read those three sentences together and DuckAssistBot becomes a clean test instrument. Refusing it costs a publisher no ranking, no search presence and no training exposure it did not already have, and buys exactly one thing: removal from a surface that cites its sources. So the rate at which sites refuse it is close to a direct reading of how many of them have decided they do not want to be cited, uncontaminated by the fear of losing search. This run measured that rate. Lantad requested https://host/robots.txt once for each of the 1,419 hostnames in this repository's two committed corpus seed files, as LantadBot, redirects followed, from one network location, and parsed every file that came back with the matcher this scanner ships. No DuckAssistBot request was observed and none could be: everything below is what 1,063 files say, not what any crawler did. The token is also absent from the fifteen in this scanner's own crawler registry, which is the first thing this run found and the first thing it should say, because a scanner that reports on AI visibility does not get to count an omission in somebody else's file without naming its own.
In short
- DuckAssistBot was named in 63 of the 1,063 robots.txt files Lantad parsed on 7 October 2026, and only 15 of those 63 files closed the site root to it: naming a crawler token and ruling on it are different acts, and in this corpus the first happens four times as often as the second.
- DuckDuckGo's own help page for the crawler, opened on 7 October 2026 and carrying no date, states that DuckAssistBot crawls pages in real-time for our AI-assisted answers, which prominently cite their sources, that this data is not used in any way to train AI models, and that opting out of DuckAssistBot does not impact organic search rankings and will not affect whether or not websites appear in our search results.
- 26 of the 1,063 files closed the root to DuckAssistBot, and 11 of the 26 never mentioned the token: those 11 caught it with a wildcard group that refuses every unnamed fetcher, which means a refusal rather than a decision about this crawler.
- 39 of the 63 files that named DuckAssistBot put it in one group of 30 user agent tokens carrying 28 path rules and no root restriction, and 35 of those 39 sites are in this corpus's Wix and Squarespace stratum, so the single commonest reason the token appears in a robots.txt file is a platform default rather than an editorial choice.
- DuckDuckGo publishes the crawler's addresses at duckduckgo.com/duckassistbot.json and its search crawler's at duckduckgo.com/duckduckbot.json, and on 7 October 2026 both files held the same 486 IPv4 prefixes, every one a /32, under the same creationTime of 2026-09-01T12:44:58, so the user agent string is the only published way to tell the two apart.
| Stage | Files | What happened |
|---|---|---|
| Hostnames asked | 1,419 | 392 in ten platform strata, 1,027 in eight industry strata |
| Never returned a status | 30 | 20 failed to resolve, 9 hit the timeout, one failed in the client |
| Answered HTTP 403 | 162 | Refused this crawler the file |
| Answered HTTP 503 | 52 | Served nothing to this client |
| Answered HTTP 404 | 51 | No file, which allows every crawler |
| Answered another status | 18 | 7 of 202, 4 of 429, 2 of 401, one each of 406, 418, 451, 498 and 529 |
| Answered HTTP 200 | 1,106 | Of which 34 returned an HTML page, not a robots.txt file |
| Parsed as plain text | 1,063 | The denominator for every rate below |
| Named DuckAssistBot | 63 | Of 210 files naming any of 16 AI crawler tokens |
| Closed the root to DuckAssistBot | 26 | 15 by name, 11 through a wildcard group |
What is DuckAssistBot, and what does blocking it in robots.txt cost?
DuckDuckGo runs two documented crawlers and keeps a separate help page for each. The page for DuckAssistBot describes it as "a web crawler for DuckDuckGo Search that crawls pages in real-time for our AI-assisted answers", gives its user agent as DuckAssistBot/1.2 followed by a link to duckduckgo.com/duckassistbot.html, and tells publishers to "opt out of being a potential source for AI-assisted answers by modifying the robots.txt file" for their domains. The page for the older crawler, at duckduckgo.com/duckduckgo-help-pages/results/duckduckbot/, describes DuckDuckBot as "a web crawler for DuckDuckGo" whose job is to improve search results, gives its user agent as DuckDuckBot/1.1, and adds one line DuckAssistBot's page does not have: "It respects WWW::RobotRules", naming a specific Perl library rather than a specification.
That asymmetry is worth pausing on, because it is the kind of detail a scanner has to read carefully rather than summarise. The search crawler's page names the parser it obeys. The AI crawler's page names a mechanism, robots.txt, and an effect, that the change "will take effect after 72 hours and DuckAssistBot will stop crawling your site", without naming which robots.txt semantics it implements. Sites are therefore invited to use a file whose dialect the vendor has not specified, and the 72 hour figure is a published lag of the sort this blog has measured from the other direction when it looked at how long a robots.txt edit takes to reach a crawler. Both pages are thin by the standards of a specification and generous by the standards of this field, where six of the nine vendors we track publish exactly one crawler token.
The practical shape of the opt out is set by RFC 9309, whose section 2.2.1, The User-Agent Line, requires that crawlers use case-insensitive matching to find the group that matches the product token and then obey the rules of that group. There is no header, no meta tag and no dashboard in DuckDuckGo's instructions. One string in one file is the entire control surface. That makes DuckAssistBot the cleanest member of a class this blog keeps separating from training crawlers: the citation fetcher, which exists to produce an answer with a link in it. The split matters because the two are refused at very different rates, which is what the counting below is for, and which is the same division that produced the gap between GPTBot in 82 of 718 files and its search twin in 24.
| Property | DuckAssistBot | DuckDuckBot |
|---|---|---|
| Stated purpose | Real time fetch for AI-assisted answers | Improving search results |
| User agent | DuckAssistBot/1.2 | DuckDuckBot/1.1 |
| Trains a model | Documented as not used for training | Not addressed |
| Robots parser named | None named | WWW::RobotRules |
| Opt out mechanism | Disallow the token in robots.txt | Not described on the page |
| Lag after the edit | 72 hours, stated | Not stated |
| Cost of opting out | No ranking or presence effect, stated | Not stated |
| Published IP list | duckassistbot.json | duckduckbot.json |
| In this scanner's registry | No | No |
How many robots.txt files name DuckAssistBot?
63 of the 1,063 files that parsed as plain text named DuckAssistBot in a user agent line. 26 of the 1,063 closed the site root to it. Those two numbers are not two views of one fact, and the distance between them is the finding: 48 of the 63 files that wrote the token down left the root open to the crawler anyway, while 11 of the 26 that closed the root never wrote the token at all. Naming is not deciding, and in this corpus the two acts barely overlap.
Set against the rest of the field, the refusal rate is low but not the lowest. Ranked by how many of the 1,063 files closed the root to each token, the order runs GPTBot at 84, CCBot at 83, Bytespider at 77, ClaudeBot at 76, Google-Extended at 64, Applebot-Extended at 59, PerplexityBot at 48, ChatGPT-User at 46, Perplexity-User at 28, OAI-SearchBot at 27, DuckAssistBot at 26, Claude-SearchBot at 24, DuckDuckBot at 8 and Bingbot at 6. The citation crawlers cluster at the bottom and the training crawlers at the top, which is the pattern a reader would hope for, and DuckAssistBot sits exactly where its documentation suggests it should: refused about as often as OAI-SearchBot and Claude-SearchBot, and about a third as often as GPTBot.
Four distinct user agent values containing the string duck appeared across 74 files: duckassistbot in 63, duckduckbot in 12, and one file each carrying duckduckbot_not and a bare duckduckgo. Only two files in the whole corpus named both of DuckDuckGo's published tokens, sciencedirect.com and theregister.com, and both of those turn out to have thought about the distinction rather than inherited it. The 63 files naming DuckAssistBot run from 800 bytes to 44,665 with a median of 1,514, so this is not a signal that only large or heavily maintained files carry. Checking which group your own file hands a given token to is what the robots.txt tester is for, and it is the same evaluation that found 232 of 1,059 files carrying a defect and that depends on the file arriving at all, which it does not always do: 92 of 1,056 sites refused GPTBot the file itself. The wider token census, 2,209 distinct tokens across 1,004 files, is the backdrop against which 63 is a small number.
Why 39 of the 63 namings are a platform default, not a decision
The reason 48 of the 63 files name the token and admit the crawler anyway is visible in the shape of the group the token sits in. 39 of the 63 put DuckAssistBot in a single group headed by 30 consecutive user agent lines, carrying 28 path rules and no root restriction. The 30 lines are identical from site to site: ai2bot, ai2bot-dolma, aihitbot, amazonbot, anthropic-ai, applebot-extended, bytespider, ccbot, claudebot, cohere-ai, cohere-training-data-crawler, duckassistbot, facebookbot, google-extended, googleother, googleother-image, googleother-video, gptbot, img2dataset, meta-externalagent, mycentralaiscraperbot, omgili, omgilibot, quora-bot, tiktokspider, youbot, three AdsBot-Google variants, and finally an asterisk. The 28 rules below them restrict /config, /search, /account, /commerce/digital-download/, /api/ and a handful of similar paths.
Putting the wildcard in the same group as 29 named crawlers is what makes the whole stack inert. Every one of those tokens receives precisely the rules an anonymous fetcher receives, so a file that reads like a considered AI policy functions as a default. 35 of the 39 sites with this signature are in this corpus's Wix and Squarespace stratum, three are in the no code stratum and one is a local media site, and not one of the 54 usable Wix and Squarespace files closed the root to DuckAssistBot. This is the same artefact measured head on when 61 of 66 Squarespace sites named 26 AI crawlers and blocked none, and it is the mirror image of the 69 Wix files that named no AI crawler at all and of the 79 of 86 Drupal files that named none, 68 of them shipping Drupal's own file unedited. A platform can get a token into a million files without getting a decision into any of them.
Strip the 39 platform defaults out and 24 files remain that named DuckAssistBot for a reason of their own. 15 of those gave it a group to itself, and 12 of the 15 closed the root. The other nine put it in a group of between three and 128 tokens, which is a judgement about a category rather than about this crawler, and that is how 89 of 145 files came to rule only on the whole site. So of the 63 appearances, 12 are an unambiguous refusal of this crawler by name, three are an unambiguous admission of it by name, and the remaining 48 are a token riding in a list, three of those lists carrying a blanket refusal of everything named in them.
Flow: 1,063 files parsed to 63 name the token; 1,063 files parsed to 1,000 do not name it; 63 name the token to 15 give it its own group; 63 name the token to 48 in a token stack; 1,000 do not name it to Falls to the wildcard group; 15 give it its own group (12) to 26 close the root; 15 give it its own group (3) to 1,037 leave the root open; 48 in a token stack (45) to 1,037 leave the root open; 48 in a token stack (3) to 26 close the root; Falls to the wildcard group (11) to 26 close the root; Falls to the wildcard group (989) to 1,037 leave the root open.
The 26 files that closed the root, and the three that got the distinction right
13 of the 26 refusals came from this corpus's news stratum, which holds 66 usable files. That concentration is not a surprise and it is consistent with what this blog found when 77 of 706 files blocked a citation crawler and 42 of those were news sites. What is more interesting is the mechanism. 11 of the 26 never named DuckAssistBot: their wildcard group carries a Disallow of the site root, so the crawler is refused as one of the anonymous many rather than as itself, which is the same effect that closed 559 of the 581 pages GPTBot lost through a rule that never named it. Two of those 11, sap.com and gmarket.co.kr, named GPTBot in a group of its own that restricts nothing at the root and left DuckAssistBot to the wildcard, which is a file that has considered AI crawler access and still refuses the one whose refusal was documented as free.
18 of the 26 still leave DuckDuckBot a path to the root while refusing DuckAssistBot, and eight refuse both. The 18 are the configuration DuckDuckGo's page describes: out of the answers, in the index. theregister.com is the clearest instance in the corpus. Its file runs to 68 groups, grants DuckDuckBot an explicit Allow of the root near the top, and 200 lines later gives DuckAssistBot a group of its own carrying a single Disallow of the root under a comment reading "DuckDuckGo AI assistant (distinct from DuckDuckBot search)". Somebody read both help pages. The same deliberate split, pointed the other way, is in sciencedirect.com, which groups Baiduspider, DuckDuckBot and DuckAssistBot together under 69 path rules that leave the root open, with a comment saying the crawler "powers DuckDuckGo's AI-assisted answers feature" and that "it is allowed here as a search companion to DuckDuckBot, not as a training crawler". Those two files disagree about the right answer and agree about the question, which is the mark of a decision.
The two near misses are worth more than the 26 refusals, because they are the failure mode the documentation invites. instacart.com writes a heading reading "Rules for DuckDuckGo Search" and under it a group for the token DuckDuckGo, with an Allow of the root and five careful path exclusions. Under the matching rule in Google's robots.txt documentation, Last updated 2026-08-31 UTC, a crawler takes "the group with the most specific user agent that matches the crawler's user agent" and every other group is ignored. DuckDuckGo publishes no token named DuckDuckGo, and that string is not a prefix of either DuckDuckBot or DuckAssistBot, so neither crawler selects the group written for it and both fall through to a wildcard group whose Disallow covers the root. A site that intended to welcome DuckDuckGo refuses both its crawlers. amsterdam.nl fails in the opposite direction: it uses a _NOT suffix as a homemade disable switch, applying it to BingBot_NOT and DuckDuckBot_NOT, so the permissive DuckDuckBot_NOT group with its Dutch comment identifying the DuckDuckGo search engine matches nothing, while a live DuckDuckBot group above it disallows the root. Renaming a token to turn a group off works, and it leaves the file saying something it no longer does, which is the same stale group problem as a crawler token that gets renamed under you.
| Host | Stratum | Verdict came from | DuckDuckBot | GPTBot |
|---|---|---|---|---|
| bloomberg.com | news | A group naming it alone | Open | Closed |
| cnn.com | news | A group of 77 tokens | Open | Closed |
| theregister.com | news | A group naming it alone | Open | Closed |
| theglobeandmail.com | news | A group naming it alone | Open | Closed |
| nzherald.co.nz | news | A group naming it alone | Open | Closed |
| france24.com | news | A group naming it alone | Open | Closed |
| dr.dk | news | A group naming it alone | Open | Closed |
| nos.nl | news | A group naming it alone | Open | Closed |
| ansa.it | news | A group naming it alone | Open | Closed |
| news24.com | news | A group naming it alone | Open | Closed |
| rappler.com | news | A group naming it alone | Open | Closed |
| sfchronicle.com | news | A group of 22 tokens | Open | Open |
| eluniversal.com.mx | news | The wildcard group | Open | Closed |
| qcitymetro.com | media-local | A group naming it alone | Open | Closed |
| amazon.com | ecommerce | A group naming it alone | Open | Closed |
| instacart.com | ecommerce | The wildcard group | Closed | Closed |
| gmarket.co.kr | ecommerce | The wildcard group | Closed | Open |
| congress.gov | government | A group of 128 tokens | Open | Closed |
| amsterdam.nl | government | The wildcard group | Closed | Closed |
| wa.gov | government | The wildcard group | Closed | Closed |
| helsinki.fi | education | The wildcard group | Closed | Closed |
| scielo.org | education | The wildcard group | Closed | Closed |
| pennmedicine.org | healthcare | The wildcard group | Closed | Closed |
| sap.com | saas | The wildcard group | Open | Open |
| botcity.dev | saas-marketing | The wildcard group | Open | Closed |
| coralvilleanimalhospital.com | wordpress-smb | The wildcard group | Closed | Closed |
One IP list for two crawlers, and what this run did not measure
There is a second control surface a publisher might reach for, and on this vendor it does not exist. DuckDuckGo publishes the crawler's addresses as JSON at duckduckgo.com/duckassistbot.json and the search crawler's at duckduckgo.com/duckduckbot.json. Both were fetched for this post on 7 October 2026. Each holds 486 IPv4 prefixes, every one a /32 host address, and the two sets are identical: 486 shared, none unique to either file. Both carry the same creationTime value of 2026-09-01T12:44:58. The AI answer crawler and the search crawler arrive from the same addresses, so no firewall rule, rate limit or edge policy keyed on IP can separate the surface a publisher wants from the one it does not. The product token in robots.txt is not merely the documented lever. It is the only one.
That is a sharper version of a pattern this blog has recorded elsewhere, and it is worth saying that the pattern is the norm rather than a lapse. Anthropic publishes three named bots behind one IP list that cannot tell them apart, while OpenAI's crawler documentation, which carried no date when it was opened on 7 October 2026, is the exception: it lists four crawlers and publishes a separate address file for each one, at openai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json and openai.com/adsbot.json. A list of addresses answers the question of whether a request is genuine. It was never built to answer the question of what the request is for, and treating it as an access control is a category error that happens to be cheap to make.
Four things this run did not establish, because the gap between them and what it did establish is where a measurement like this normally goes wrong. It did not observe DuckAssistBot. Every request was sent by this scanner as LantadBot from one network location, so nothing here is evidence that DuckDuckGo honoured any rule, respected the 72 hour window it documents, or fetched any page in this corpus. It did not test any path but the site root, so a file that admits the crawler at the root while closing an articles directory counts as open here. It did not reach the 356 hostnames that returned no file this scanner could parse, and the 162 that answered the request with an HTTP 403 are exactly the sites most likely to hold firm views about crawler access, so the surviving 1,063 lean towards the permissive and every refusal rate above is probably a floor. And the corpus is an editorial sampling frame built for platform and industry coverage rather than a random draw of the web, so each figure describes these 1,419 hostnames on one date and nothing wider. The scoring rules this scanner applies to what it does measure are set out on the methodology page, and DuckAssistBot is not among the tokens they cover, which is a gap in our registry and now a logged one.
duckassistbot.json
- 486 prefixes, all IPv4
- Every entry a /32 host address
- creationTime 2026-09-01T12:44:58
- 0 addresses unique to this file
duckduckbot.json
- 486 prefixes, all IPv4
- Every entry a /32 host address
- creationTime 2026-09-01T12:44:58
- 0 addresses unique to this file
Lantad
Published .
Most of what this blog counts about AI crawler access is a question with an uncomfortable answer: a site blocks a training crawler and loses a citation crawler by accident, or blocks a citation crawler and loses search traffic it wanted. DuckAssistBot is the rare case where the vendor has removed the ambiguity in writing. DuckDuckGo publishes a help page for the crawler at duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/, which was opened and read for this post on 7 October 2026 and carries no publication date. It says the crawler "crawls pages in real-time for our AI-assisted answers, which prominently cite their sources", that "this data is not used in any way to train AI models", and that "opting out of DuckAssistBot does not impact organic search rankings and will not affect whether or not websites appear in our search results".
Common questions
What is DuckAssistBot?
DuckAssistBot is DuckDuckGo's crawler for AI-assisted answers. Its own help page, read on 7 October 2026, says it crawls pages in real-time for our AI-assisted answers, which prominently cite their sources, and that this data is not used in any way to train AI models. Its user agent is DuckAssistBot/1.2 followed by a link to duckduckgo.com/duckassistbot.html. It is a separate crawler from DuckDuckBot, which the vendor describes as a web crawler for DuckDuckGo whose job is improving search results.
Does blocking DuckAssistBot in robots.txt hurt search rankings?
DuckDuckGo says it does not. The help page for the crawler states that opting out of DuckAssistBot does not impact organic search rankings and will not affect whether or not websites appear in our search results. That is the vendor's claim about its own product and Lantad has not tested it. What Lantad did measure is that 18 of the 26 files in this corpus that closed the root to DuckAssistBot left DuckDuckBot a path to it, which is the configuration that claim describes.
How long does a DuckAssistBot robots.txt change take to apply?
72 hours, according to DuckDuckGo. Its help page states that once the file is updated to disallow the DuckAssistBot user agent, the change will take effect after 72 hours and DuckAssistBot will stop crawling your site. The page does not say which robots.txt parser the crawler uses, unlike the page for DuckDuckBot, which says it respects WWW::RobotRules.
Why does my robots.txt mention DuckAssistBot when I never added it?
Most likely because your platform added it. 39 of the 63 files in this corpus that named DuckAssistBot on 7 October 2026 placed it in one identical group of 30 user agent tokens with 28 path rules and no root restriction, and 35 of those 39 sites are built on Wix or Squarespace. That group also contains the wildcard, so every token in it receives exactly the rules an unnamed crawler receives. The token being present is not evidence that anything is blocked.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.