BlogFindings
Do AI crawlers read PDFs? 332 on 124 home pages, and 297 carried no X-Robots-Tag
Lantad requested the home page of all 1,419 hostnames in this repository's committed corpus on 24 September 2026 and read the raw bytes with no JavaScript executed. 1,085 answered HTTP 200 with HTML, and 124 of those linked a document file. 332 of the 335 document links were PDFs, 317 of the fetched files really carried PDF bytes, and 297 of those 317 sent no X-Robots-Tag, which is the only control Google documents for a file that is not HTML.
This run went looking for the supply side of that question rather than the demand side. Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for their robots.txt and their home page on 24 September 2026, read every home page as raw bytes with no JavaScript executed, collected every link that ended in one of thirteen document extensions, and then requested each of those documents and recorded what came back. Nothing here is a measurement of any AI crawler: every request came from this scanner's own token. What it measures is what a crawler would find if it went, which is the half of the problem a site owner can actually change.
In short
- Lantad requested the home page of all 1,419 hostnames in this repository's committed corpus on 24 September 2026 and read the raw bytes with no JavaScript executed: 1,085 answered HTTP 200 with HTML, 124 of those linked at least one document file, and 332 of the 335 document links were PDFs.
- Whether AI crawlers read PDFs is documented by none of the companies that operate one: the crawler pages published by OpenAI, Anthropic and Perplexity name no file format, while Google's file types documentation carrying Last updated 2026-02-03 UTC lists twenty document formats it can index, and only four of those twenty appeared anywhere in this corpus.
- Google's robots meta tag documentation, carrying Last updated 2026-03-24 UTC, states that the X-Robots-Tag response header is the way to control indexing of a non-HTML resource such as a PDF, and 297 of the 317 PDFs Lantad fetched on 24 September 2026 carried no such header.
- Of the 191 same-host document links Lantad could evaluate against a parsed robots.txt on 24 September 2026, 174 were allowed to GPTBot and 17 were disallowed, and 11 of those 17 were decided by a wildcard group that never names GPTBot.
- Google's Googlebot documentation, carrying Last updated 2026-02-03 UTC, states that Googlebot crawls the first 64MB of a PDF file, and one of the 317 PDFs Lantad fetched on 24 September 2026 exceeded that ceiling, at 69,408,115 bytes.
| Format | On Google's indexable list | Links found in this corpus | Fetched and confirmed |
|---|---|---|---|
| Adobe Portable Document Format | Yes | 332 | 317 |
| Microsoft PowerPoint | Yes | 1 | 1 |
| Microsoft Excel | Yes | 1 | 1 |
| Microsoft Word | Yes | 1 | 1 |
| PostScript, EPUB, Hanword, RTF, OpenOffice | Yes | 0 | 0 |
| CSV, KML, GPX, TeX, WML | Yes | 0 | 0 |
Do AI crawlers read PDFs, and has anyone said so?
The honest answer is that nobody who runs one has written it down. Google's file types documentation, carrying Last updated 2026-02-03 UTC, names nine flat file types and eleven encoded document types it can index, puts Adobe Portable Document Format at the head of the encoded list, and states that the file type is determined by the Content-Type HTTP header returned when Google crawls the file, with the extension used as a fallback when that header is missing or incorrect. That is a complete answer for one crawler, published, dated and revisable.
For the crawlers that feed AI answers there is no equivalent. This site established that in August, when three AI crawler documentation pages were read and none of them named a single file type: the strings PDF, file type and Content-Type appeared zero times across all three. OpenAI's crawler documentation describes what GPTBot, OAI-SearchBot and ChatGPT-User are for and how to block each one, and says nothing at all about what any of them does when the thing at the end of the URL is a 4MB annual report. That silence is not evidence that those crawlers cannot parse a PDF. It is evidence that a site owner cannot find out, which is a different problem and a more useful one to name, because it tells you exactly where the guessing starts.
The pattern is consistent rather than a one-off gap. Vendor crawler pages in this category are written to answer one question, which is how to block the crawler, and they answer it well. They do not describe operating behaviour: no crawler vendor documents Retry-After either, and when this site counted the published tokens, six of the nine vendors tracked published exactly one crawler token, which is not enough to tell a training fetch from a search fetch. Read against that background, the absence of a file format list is ordinary, and planning around it as though the crawlers behave like Googlebot is a guess wearing a citation.
So this run did the part that can be established. It did not ask what GPTBot does with a PDF. It asked how many sites put one in a crawler's path at all, whether their own robots.txt lets a crawler have it, and what the response headers say when the file is requested. Every one of those is a fact about the site, recorded on a date, and every one of them is something the site owner controls.
-
Google file types pageTwenty formats named Content-Type decides the parser, extension is the fallback, Last updated 2026-02-03 UTC -
Google robots meta tag pageNames the PDF control X-Robots-Tag is the way to control a non-HTML resource, Last updated 2026-03-24 UTC -
OpenAI crawler documentationNo format named Three agents described, robots.txt tokens listed, no date on the page -
Anthropic and Perplexity crawler pagesNo format named Read 19 August 2026 and published here: zero occurrences of PDF, file type or Content-Type
How many home pages put a document in front of a crawler
Of the 1,419 hostnames asked, 1,085 answered the home page request with HTTP 200 and an HTML content type, and those 1,085 pages are the denominator for everything below. Across them 189,060 link targets were resolved, and 335 of those targets ended in a document extension. They were spread across 124 pages, so 961 of the 1,085 home pages offered a crawler no document at all.
Two things about the 335 are worth more than the headline. The first is the format concentration. 332 were PDFs. The other three were one PowerPoint deck, one Excel workbook and one Word document, one each, on three different sites. Sixteen of Google's twenty documented document formats appeared zero times: not one PostScript file, EPUB, Rich Text Format file, OpenOffice document or CSV was linked from any home page in the corpus. A list of twenty formats describes what one crawler can parse. In practice the web put one format in front of it, and any plan that treats document handling as a multi-format problem is solving a problem this corpus does not have.
The second is where the documents were. 216 of the 335 links pointed at the same host as the page carrying them and 119 pointed somewhere else, across 45 distinct external hostnames, most of them a content delivery network or a digital asset management subdomain belonging to the same organisation. That split matters because a robots.txt governs one host and nothing else, so for 119 of these links the rules that decide access were never in the file the site owner edits.
The distribution by sector is not even, and it is not random either. Of the 94 finance home pages that answered, 32 linked a document, and finance carried 116 of the 335 links on its own. Healthcare followed at 22 of 96, travel at 16 of 79, education at 15 of 102, government at 9 of 87 and SaaS at 10 of 119. At the other end, none of the 44 single page app startups and none of the 34 Shopify direct to consumer stores linked a document from the home page at all. The sites carrying the regulated disclosure are the sites carrying the PDFs, which is the expected result and is worth stating precisely because it locates who this finding is about.
This is the same shape found every time this corpus has been read for a particular element. 87 table elements appeared on 1,079 home pages and two of them carried a caption; 52,077 images carried 18,171 empty alt attributes; 834 iframes on 433 pages held 9,098 words between them. A home page is a poor place to look for depth, and a document link is one of the few signals on it that points at something a company had to be careful about.
What robots.txt actually does to a PDF link
A robots.txt rule matches a path, and it does not care what is at the end of it, so a PDF is governed by exactly the same machinery as a page. 191 of the 216 same-host document links sat on a host whose robots.txt returned a plain-text file that parsed into at least one group, and each of those 191 paths was evaluated for GPTBot with parseRobotsTxt and evaluateRobots from core/src/robots.ts, so the verdict in this post is the verdict this product's own tester would give.
174 of the 191 were allowed. 17 were disallowed, across 9 hostnames. The interesting number is the split underneath: 6 of the 17 were decided by a group that names GPTBot, and 11 were decided by the wildcard group, meaning the rule that closed the document was written without any crawler in mind. Group selection working that way is RFC 9309 behaving as specified rather than any file misbehaving, and it is the dominant mechanism at every scale this corpus has been read at: when 58,813 linked paths were evaluated last week, 559 of the 581 that GPTBot lost were closed by a rule that never named it.
The wildcard cases here have a recognisable shape. Six of the 11 are one bank whose robots.txt disallows a digital asset management directory, which is where its PDFs live, so a rule written to keep a crawler out of an asset tree took the rate cards with it. Two more are a cruise line refusing a travel documents directory. None of that looks like a decision about AI at all, and that is the point: nobody blocked a PDF, somebody blocked a directory, and a PDF was in it. The only way to find out is to evaluate the actual path rather than read the file as prose, which is what the robots.txt tester on this site does per token.
Two limits on this section, both of which narrow it. The first is that 25 of the 216 same-host links could not be evaluated because their host served no usable robots.txt, and 15 of the 124 document-linking hosts fell into that group. The second, and larger, is that 119 of the 335 links point at another host entirely, and a crawler fetching one of those reads that host's robots.txt, not the site's. Of the 124 hosts linking a document, only 12 named any of the 15 AI crawler tokens in this scanner's registry anywhere in their robots.txt. For the other 112 the question of whether an AI crawler may take the annual report has never been answered by the site, in either direction.
None of this says anything about whether a crawler obeyed the file. A robots.txt is a declaration, not an observation, and 92 of 1,056 sites refused GPTBot's user agent string the robots.txt file itself when that was measured on 18 September 2026. The declaration and the server's behaviour are separate measurements and this one is the declaration.
| Verdict for GPTBot | Document links | Hostnames | Decided by |
|---|---|---|---|
| Allowed | 174 | Across the evaluable hosts | No matching rule, or an allow rule |
| Disallowed by a group naming GPTBot | 6 | 4 | A deliberate, named rule |
| Disallowed by the wildcard group | 11 | 5 | A rule written without any crawler in mind |
| Not evaluable | 25 | 15 of the 124 document hosts | No plain-text robots.txt returned |
| Governed by another host's file | 119 | 45 external hostnames | A CDN or asset domain, not the site's own file |
The only header Google documents for a PDF, and the 297 that did not send it
Every one of the 335 document links was then requested. 324 answered HTTP 200, five answered 404, three answered 403, two answered 503 and one timed out. Of the 321 PDF links that answered 200, 317 really carried PDF bytes, and those 317 came from 117 distinct hostnames. They are the population for the rest of this post.
Here is where the interesting gap sits. A PDF has no head element, so it cannot carry a robots meta tag, and Google's robots meta tag documentation, carrying Last updated 2026-03-24 UTC, says so directly: to block indexing of non-HTML resources such as PDF files, video files or image files, use the X-Robots-Tag response header instead. That header is therefore the entire published control surface for a PDF. 20 of the 317 carried one, from 8 hostnames. 297 carried nothing.
Of those 20, 18 restrict indexing and 2 do not. Four banks send noindex on their PDFs and two smaller sites send noindex, nofollow, which taken together is the only group in this sample that appears to have made a decision about documents at all. One footwear retailer sends none, the strongest value available. Two files on one SaaS site send the value all, which is the default and changes nothing, so it is a header that was configured rather than a control that was applied. This site has read that header across the corpus before in its HTML form, where 5 of 391 home pages sent an X-Robots-Tag and all five turned out to be a captcha, so a rate of 20 in 317 on documents is, oddly, the higher adoption of the two.
The honest framing of the 297 is not that they are misconfigured. A missing X-Robots-Tag means indexable, which is very probably what those organisations want for a brochure or a product sheet, and treating silence as a defect would be exactly the inference this site refuses to make elsewhere. What the number establishes is narrower and harder to argue with: on 297 of 317 documents, no instruction of any kind was attached to the file, so whatever an AI crawler does with those bytes, it will not be because the site asked. And the instruction that is usually assumed to cover this does not: noindex is an indexing directive, not an access one, and only 2 of 9 operator pages name the tag at all when that was checked.
| X-Robots-Tag value returned | PDFs | Hostnames | What it asks for |
|---|---|---|---|
| No header sent | 297 | 111 | Nothing, which resolves to indexable |
| noindex | 9 | 4 | Keep the document out of the index |
| noindex, nofollow | 6 | 2 | Keep it out and do not follow its links |
| none | 3 | 1 | Equivalent to noindex, nofollow |
| all | 2 | 1 | The default, so no restriction at all |
What the bytes said about content type, size and age
Google's file types page states that the Content-Type header decides which parser runs, which makes the header a functional part of whether a document is readable rather than a formality. 315 of the 324 documents that answered 200 sent a Content-Type matching the extension in their URL. Four sent application/octet-stream, all four from one bank, and one sent no Content-Type header at all. Google's page describes a fallback to the extension for exactly this case, but no AI crawler vendor documents having one.
Four responses were a different problem. Four URLs ending in .pdf on one central bank's site answered HTTP 200 with a text/html content type and bodies between 44,233 and 50,145 bytes. There is no PDF at those addresses: a crawler following the link receives an HTML page, and so does a person. That is a broken link that returns a success code, which is the same failure mode measured across this corpus as a soft 404, and it is invisible to any check that only looks at status codes.
On size, Google's Googlebot documentation, carrying Last updated 2026-02-03 UTC, states that Googlebot crawls the first 2MB of a supported file type and the first 64MB of a PDF file, and that the limit is applied to uncompressed data. The corpus sits well inside that. The median PDF was 408,990 bytes, 66 of the 317 exceeded 2MB, 17 exceeded 10MB and exactly one exceeded the 64MB PDF ceiling, a 69,408,115 byte file on a medical institute's site. The general 2MB limit is the one that bites more often, and this site has covered it before: Googlebot reads the first 2MB and three AI crawler documentation pages name no limit at all, so for 66 of these documents there is a published ceiling from one vendor and silence from the rest.
On freshness, 290 of the 317 sent a Last-Modified header and 247 sent an ETag, which is markedly better than the corpus manages on its HTML: when change signals were measured across home pages, 250 of 647 sent no signal that anything had changed. The dates themselves are older than the pages that link them. The median document was last modified 190 days before the scan and 30 were more than three years old, the oldest by 8.9 years. A PDF is the part of a site nobody revisits, and a crawler asked to summarise a company's current position will find last decade's fee schedule sitting at a live URL with a 200 beside it.
Two further counts are reported with their limits attached rather than as findings. Nine PDFs on five hostnames carried an /Encrypt dictionary, which means permissions are set on the file; this run did not attempt to open any of them, so nothing is claimed about whether text extraction succeeds. And 39 PDFs on 21 hostnames carried no /Font token anywhere in the delivered bytes. That is often the signature of a scanned page image with no text layer, but it is not proof of one, because a PDF that puts its objects in compressed object streams can hide the token from a byte search, and this scanner did not decompress them. The number is a prompt to check those files, not a verdict on them, and stating it that way is the same discipline behind refusing to grade a page that could not be measured.
What to check on your own site this week
Nothing above tells you what GPTBot or ClaudeBot does with a PDF, and no honest reading of it can, because the vendors have not said and this run did not ask them. What it does give you is a short list of things that are true of your own site, that you control, and that you can settle in an afternoon. The methodology page describes how this scanner separates what was measured from what was inferred, and this list stays firmly on the measured side.
Start with the path, not the file. Take the document URLs your home page and your main navigation actually link, and evaluate each full path against your robots.txt for a named crawler token rather than reading the file and forming an impression. 11 of the 17 disallowed documents in this corpus were closed by a wildcard rule aimed at a directory, and not one of those sites will have thought of itself as blocking a PDF. If the document sits on a CDN or an asset subdomain, the file that governs it is that host's, which is the case for 119 of the 335 links measured here. Seeing the fetch as a crawler receives it is the fastest way to catch the difference between what you wrote and what applies.
Then check the response, because two of this run's findings are invisible from the browser. Confirm the Content-Type is right for the format, since that header decides which parser runs and a document served as application/octet-stream is relying on a fallback that only one vendor documents. Confirm the bytes are actually the format the extension promises: four .pdf URLs in this corpus returned an HTML page with a 200, which no link checker that stops at the status code will ever report. And decide, once, whether you want an X-Robots-Tag on documents at all. 297 of 317 sent nothing, and the defensible position is either a deliberate nothing or a deliberate directive, not an accident.
Finally, treat the document like content rather than like an artefact. The median file here was last modified more than six months before the page linking it and thirty were over three years old, which means a crawler can quote a superseded price from a live URL and be accurate about what it found. If a document carries a figure you would correct on the page, it needs the same review cycle as the page. And if the substance a buyer needs is only in the PDF, the strongest move is not a crawler setting: it is putting the same facts in HTML on a page, where structured data can label them and where every crawler in the category has documented that it can read them. This scanner's own token and the rules it follows are published at the crawler policy page, which is the standard the vendors in this post have not met.
- Evaluate the document path, not the robots.txt file 11 of 17 disallowed documents were closed by a wildcard rule aimed at a directory
- Check which host governs the document 119 of 335 links pointed at a CDN or asset domain with its own robots.txt
- Confirm the Content-Type matches the format 5 of 324 responses sent octet-stream or no type at all
- Confirm the bytes match the extension 4 .pdf URLs returned an HTML page with HTTP 200
- Decide on X-Robots-Tag deliberately 297 of 317 sent no header, which resolves to indexable by default
- Check the document against Google's 64MB PDF ceiling One file exceeded it at 69,408,115 bytes; 66 exceeded the 2MB general limit
- Review the document's date as you would the page's Median last modified 190 days earlier, 30 files older than three years
Lantad
Published .
Most of what a bank, a hospital or a university actually commits to writing does not live on a web page. It lives in the fee schedule, the annual report, the patient information leaflet, the admissions handbook. Those documents sit at real URLs, they are linked from the home page, and they are almost always PDFs. If an answer engine is going to quote a company's own terms back to a customer, this is where the terms are. So the question of whether an AI crawler can get at them is not a filing detail, and it has never had a published answer.
Common questions
Do AI crawlers read PDFs?
No vendor that operates one has published an answer. The crawler documentation from OpenAI, Anthropic and Perplexity names no file format, no Content-Type behaviour and no size limit, which was established by reading all three pages on 19 August 2026. Google is the exception and documents that it indexes twenty document formats including PDF, with the Content-Type header deciding the parser. So for the crawlers that feed AI answers the behaviour is unspecified rather than known, and the honest position is that it is a guess.
How do I stop an AI crawler reading a PDF on my site?
A PDF has no head element, so a robots meta tag cannot be placed on it. Google's robots meta tag documentation, carrying Last updated 2026-03-24 UTC, states that the X-Robots-Tag response header is the way to control indexing of a non-HTML resource. A robots.txt Disallow covering the document's path also applies normally, since a robots.txt rule matches a path and does not care about the format. Of 317 PDFs measured on 24 September 2026, 297 carried no X-Robots-Tag at all.
Does a robots.txt rule apply to a PDF the same way it applies to a page?
Yes. A robots.txt rule matches a path, so a directory-level Disallow closes every document inside it as readily as every page. That is how most blocked documents in this corpus were blocked: of 17 document paths disallowed to GPTBot on 24 September 2026, 11 were decided by the wildcard group and only 6 by a group naming GPTBot. A document hosted on a CDN or asset subdomain is governed by that host's robots.txt instead, which applied to 119 of the 335 document links measured.
How large can a PDF be before a crawler stops reading it?
Google publishes a figure and the AI crawler vendors do not. Google's Googlebot documentation, carrying Last updated 2026-02-03 UTC, states that Googlebot crawls the first 2MB of a supported file type and the first 64MB of a PDF file, applied to uncompressed data. Of 317 PDFs Lantad fetched on 24 September 2026, 66 exceeded 2MB, 17 exceeded 10MB and one exceeded 64MB at 69,408,115 bytes. No comparable limit is documented by OpenAI, Anthropic or Perplexity.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.