Blog / The EU AI Act never writes robots.txt, and its Code of Practice does twice

The EU AI Act never writes robots.txt, and its Code of Practice does twice

Regulation (EU) 2024/1689 obliges general-purpose AI model providers to identify and comply with machine-readable rights reservations, and the phrase robots.txt appears nowhere in it. The Commission's Code of Practice names the file and the RFC that defines it, and the Commission's power to fine under Article 101 applies from 2 August 2026.

In short

  • Regulation (EU) 2024/1689, the EU AI Act, contains the string robots.txt zero times. Its Article 53(1)(c) instead requires providers of general-purpose AI models to put in place a policy to identify and comply with, including through state-of-the-art technologies, a reservation of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790.
  • Article 4(3) of Directive (EU) 2019/790 is where the words machine-readable come from: the text and data mining exception applies only where the rightholder has not expressly reserved the use in an appropriate manner, such as machine-readable means in the case of content made publicly available online.
  • The Copyright Chapter of the Commission's General-Purpose AI Code of Practice, published 10 July 2025, is the document that names the file: Measure 1.3 commits signatories to employ web crawlers that read and follow the Robot Exclusion Protocol as specified in IETF RFC 9309.
  • Article 113 of the AI Act states that the Regulation applies from 2 August 2026, and that Chapter V applies from 2 August 2025 with the exception of Article 101, so the Commission's power to fine a model provider up to 3 percent of worldwide turnover or EUR 15 000 000 arrives on the later date.
  • None of this is observable from outside a model provider. An external scan can read what a site publishes, and it cannot see whether any crawler obeyed it, so Lantad reports the first and says nothing about the second.

A rule that tells a machine to read something has to say what to read. The EU AI Act does not. Searched as published text, Regulation (EU) 2024/1689 contains the string robots.txt zero times, and its single occurrence of the word robots is a recital about autonomous robots in manufacturing and care. What it says instead is that a provider of a general-purpose AI model must identify and comply with a reservation of rights expressed pursuant to another instrument entirely, and it leaves the format of that reservation to be settled somewhere else.

That is a chain of three documents, and the file every site owner actually edits only appears in the third one. This post follows the chain in order: what the Regulation obliges, what the Directive it points at means by machine-readable, and what the Commission's Code of Practice commits its signatories to reading. Then it sets out the date that changes, and what an external scanner can honestly tell you about any of it.

Lantad has measured nothing here. This is a reading of published legal texts and one Commission document, quoted with links, and it is reported rather than tested. What Lantad can add is the last part: the difference between a rule that a crawler is asked to obey and a fact an external scan can observe, which is exactly the gap this site exists to be precise about.

  • Regulation (EU) 2024/1689 robots.txt: 0 Article 53(1)(c) requires a policy to identify and comply with a reservation of rights, naming no format at all.
  • Directive (EU) 2019/790 robots.txt: 0 Article 4(3) supplies the standard: reserved in an appropriate manner, such as machine-readable means. Still no format.
  • Code of Practice, Copyright Chapter robots.txt: 2 Measure 1.3 names the Robot Exclusion Protocol and RFC 9309, and later refers to a signatory's robots.txt features.
  • What any of them measures Nothing on your site All three describe obligations on a model provider. None creates an observable property of a web page.
Three documents, and which one names the file. Term counts are from the published English texts of the two EU instruments and from the Copyright Chapter PDF of the Code of Practice, read on 29 July 2026.

What Article 53 actually obliges a model provider to do

The operative sentence is short. Article 53(1)(c) of the AI Act requires providers of general-purpose AI models to "put in place a policy to comply with Union law on copyright and related rights, and in particular to identify and comply with, including through state-of-the-art technologies, a reservation of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790".

Three things in that sentence are worth separating. The obligation is to have a policy, not to achieve an outcome. The means are described as state-of-the-art technologies, which is a moving standard rather than a named one. And the thing to be identified is defined by cross-reference, so the AI Act itself never has to decide what a reservation looks like.

The reach is wider than the drafting suggests. Recital 106 states that any provider placing a general-purpose AI model on the Union market should comply with this obligation "regardless of the jurisdiction in which the copyright-relevant acts underpinning the training of those general-purpose AI models take place", and gives the reason as a level playing field, so that no provider gains an advantage in the Union market by applying lower copyright standards. A model trained entirely outside the EU is not outside the obligation if the model is placed on the EU market.

Article 53(2) exempts free and open-source models from the documentation duties in points (a) and (b), and recital 104 says in terms that the exemption "should not concern" the training-content summary or the copyright policy. So the copyright policy obligation survives an open-source release. That matters for anyone reasoning about which AI crawler traffic is covered, because the crawler fleets behind open-weight models are not carved out. What none of this does is tell a site owner what to publish. For that the Regulation hands off, and the next document is where the word machine-readable finally appears.

  • (a) Technical documentation of the model Annex XI content, provided on request to the AI Office. Exempted for free and open-source models under Article 53(2).
  • (b) Information for downstream providers Annex XII content, for providers integrating the model. Also exempted under Article 53(2).
  • (c) A copyright policy, including rights reservations Identify and comply with a reservation expressed pursuant to Article 4(3) of Directive (EU) 2019/790. Not exempted.
  • (d) A public summary of training content Sufficiently detailed, to a template provided by the AI Office. Not exempted.
Article 53(1) of Regulation (EU) 2024/1689, its four points as published. Presence marks which points survive the free and open-source exemption in Article 53(2), read with recital 104.

Where the words machine-readable come from

The cross-reference lands on Directive (EU) 2019/790, the copyright directive that predates the current wave of models by several years. Its Article 4 creates a text and data mining exception for reproductions and extractions of lawfully accessible works. Article 4(3) then attaches the condition that decides everything: the exception "shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online".

That is the whole legal basis for the idea that a machine-readable file on your site carries weight. It is a condition on an exception, not a right to be granted, which is why the burden sits on the rightholder to express something rather than on the crawler to ask.

Recital 18 of the same Directive goes one step further and is the closest either instrument comes to naming a mechanism: for content made publicly available online, "it should only be considered appropriate to reserve those rights by the use of machine-readable means, including metadata and terms and conditions of a website or a service". Two examples, and robots.txt is neither of them. Metadata and terms and conditions are what the drafters had in mind in 2019.

This is the origin of a persistent confusion. The AI Act uses the phrase machine-readable four times in its own text, and not one of those four is about rights reservations: they concern the EU database, the CE marking, and the Article 50 duty to mark synthetic outputs as artificially generated. The machine-readable that matters for crawling is borrowed wholesale from a 2019 directive whose examples are page metadata and website terms, both of which sit outside the file most people are editing when they think about AI crawler access.

Named in recital 18

  • Metadata: attached to the asset itself
  • Terms and conditions of a website or a service
  • Both survive the content being copied off-site
  • Neither has an agreed syntax for AI training
  • Neither is fetched by a crawler before crawling

Not named in either EU instrument

  • robots.txt: one file, one host, path rules only
  • Fetched before crawling, which is why it works
  • Expresses access, not a rights reservation
  • Named only in the Code of Practice, Measure 1.3
  • Cannot travel with a copied page
The two mechanisms recital 18 of Directive (EU) 2019/790 gives as examples of machine-readable means, against the file most site owners actually change. Read from the published text, not measured.

The Code of Practice is the document that names robots.txt

The European Commission published the General-Purpose AI Code of Practice on 10 July 2025 in three chapters, Transparency, Copyright, and Safety and Security. The Copyright Chapter, authored by Working Group 1 co-chair Alexander Peukert and vice-chair Céline Castets-Renard, is a six page document, and it is the first place in this chain where the file has a name. The chapter is hosted on the Commission's digital strategy site, which is not a host this blog is permitted to link, so the URL is here as plain text: https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai

Measure 1.3, headed "Identify and comply with rights reservations when crawling the World Wide Web", commits signatories "to employ web-crawlers that read and follow instructions expressed in accordance with the Robot Exclusion Protocol (robots.txt), as specified in the Internet Engineering Task Force (IETF) Request for Comments No. 9309, and any subsequent version of this Protocol". RFC 9309 is the same Standards Track document that decides what a 404 and a 503 on robots.txt mean, and it is now load bearing in a compliance instrument as well as in a crawler.

Measure 1.3 does not stop at the file. Point (1)(b) extends the commitment to "other appropriate machine-readable protocols to express rights reservations", giving asset-based or location-based metadata as the example, and conditioning it on adoption by a standardisation body or on being state-of-the-art and widely adopted. That is a live description of a gap rather than a settled rule, and it is the same gap that leaves the IETF AI preferences vocabulary undeployable on a real site today.

Two further points deserve attention from anyone who watches crawler behaviour. Point (4) commits signatories to make public information about "the web crawlers employed, their robots.txt features and other measures" and to provide automatic notification when it changes, which is the documentation a crawler registry depends on. Point (5) encourages signatories that also run a search engine to ensure that honouring a reservation for training does not lead to adverse effects on indexing, which is the separation between training and retrieval that crawler tokens with no user agent were invented to express.

One thing the Code is not is binding. Article 53(4) of the AI Act makes codes of practice a means to demonstrate compliance until a harmonised standard is published, and the Code's own objectives state that adherence "does not constitute conclusive evidence of compliance". A provider that does not sign has to demonstrate alternative adequate means. So robots.txt is not required by law; it is the mechanism the main compliance route names.

Article 53(1)(c) resolved

  • AI Act Art 53(1)(c): identify and comply with a reservation of rights no format named
  • resolves to Directive (EU) 2019/790 Art 4(3) machine-readable means
  • Directive recital 18: examples given metadata, terms and conditions
  • AI Act Art 53(4): codes of practice may demonstrate compliance until a harmonised standard
  • Code of Practice Measure 1.3(1)(a) robots.txt, RFC 9309
  • Code of Practice objectives: adherence as evidence not conclusive
The cross-reference chain from the Regulation to the file, as written in the three documents. Each step is a quotation or a named article, not an inference.

What changes on 2 August 2026

The dates are set by Article 113, and they are worth reading literally because the headline dates in circulation flatten a distinction the text makes. Article 113 states that the Regulation "shall apply from 2 August 2026", then carves out three exceptions: Chapters I and II from 2 February 2025; "Chapter III Section 4, Chapter V, Chapter VII and Chapter XII and Article 78" from 2 August 2025, "with the exception of Article 101"; and Article 6(1) with its corresponding obligations from 2 August 2027.

Chapter V is where Article 53 lives, so the copyright policy obligation has applied since 2 August 2025. Article 101 is expressly pulled out of that early start, which leaves it on the general date. Article 101 is the enforcement provision for this specific population: it empowers the Commission to impose on providers of general-purpose AI models fines "not exceeding 3 % of their annual total worldwide turnover in the preceding financial year or EUR 15 000 000, whichever is higher", where it finds the provider intentionally or negligently infringed the relevant provisions, failed to comply with an information request, failed to comply with a requested measure, or failed to give access to the model for evaluation.

So the shape of the change is narrow and specific. The duty to have a copyright policy is a year old. What arrives on 2 August 2026 is the Commission's power to fine for failing it. That is a change in consequence rather than a change in requirement, and it is the reason the date is being written about at all.

Nothing about that date changes a byte on your website. It does not create a new file to publish, a new header to send, or a new directive syntax to adopt. If your robots.txt already expresses what you want, the same file expresses it on 3 August. What changes is on the other side of the request, inside organisations that no external scan reaches, which is the point at which honest reporting of this story has to stop and hand over to what can actually be checked. Anyone reasoning about generative engine optimization on the strength of this date should be clear that it is a regulatory milestone, not a ranking event.

  • Chapters I and II 2 February 2025 General provisions and prohibited practices. Already applying.
  • Chapter V, including Article 53 2 August 2025 The copyright policy obligation on general-purpose AI model providers. Already applying, for a year.
  • Article 101 2 August 2026 Named exception to the 2 August 2025 start, so it falls to the general date. Fines up to 3 percent of worldwide turnover or EUR 15 000 000.
  • Article 6(1) and its obligations 2 August 2027 High-risk classification rules. Not about crawling.
Application dates as set out in Article 113 of Regulation (EU) 2024/1689, quoted from the published text. Only the provisions relevant to model providers and crawling are listed.

What an external scan can see, and what it cannot

This is the part where a measurement product has to be careful, because the temptation on a regulatory news day is to imply that a scan can report compliance. It cannot, and the reason is structural rather than a limitation anyone will engineer away.

What an external scan observes is one side of a conversation. It can fetch your robots.txt and parse it, evaluate the rules that apply to a named crawler token, and report which tokens are allowed and which are disallowed. It can fetch a page as a crawler and compare what the HTML carries against what a browser renders, which is what prose parity measures. It can read structured data and check whether a machine can identify the publisher. All of that is a property of bytes your server returns, so it is measurable by anybody, repeatedly, from outside.

What it cannot observe is the other side. Whether a given model provider fetched your robots.txt, parsed it the way RFC 9309 specifies, and honoured the result is an event inside their infrastructure. A scanner sees no trace of it. Neither, in practice, do your own logs with any certainty, because a user agent is a claim rather than an identity and confirming it requires a reverse lookup against published address ranges. The industry is building toward cryptographic answers, and what Web Bot Auth actually specifies is a signature on the crawler's request rather than anything a site publishes, so even that arrives on the request path and not in a scan of your pages.

The same limit applies to the two mechanisms recital 18 names. Terms and conditions are prose on a page: readable, but there is no agreed syntax that would let any scanner say a reservation has been validly expressed rather than merely written down. Asset-level metadata sits inside image and document files rather than in the HTML most scanners parse. So a tool that reported an EU rights reservation status would be inferring it, and Lantad does not carry that field. The honest scope is narrower and still useful: what you publish, what a crawler can read, and where the two disagree.

That reticence is the same rule that makes Lantad withhold a grade when a measurement failed rather than print a confident number. A regulation whose subject is somebody else's internal policy is not a thing to score a website on.

Sample Illustrative, not a measurement of any real site.

  • robots.txt rules per crawler token Observable A file on your origin. Fetch it, parse it, evaluate it against a token. Repeatable by anyone.
  • Terms and conditions text Readable, not evaluable Prose on a page. No agreed syntax exists that would let a scanner call it a valid reservation.
  • Asset-level metadata Mostly out of reach Inside image and document files rather than in the HTML. Named in recital 18 as an example.
  • Whether a provider honoured it Not observable An event inside a model provider's crawling infrastructure. No scan of your site produces evidence either way.
  • A provider's copyright policy Not observable Article 53(1)(c) is a duty on the provider. Nothing about it is a property of your pages.
Where each element of the chain lives, and whether an external scan of a website can observe it. Illustrative of scope, not a measurement of any site.

What to check on your own site this week

None of the above obliges a site owner to do anything. The obligations in Article 53 run to model providers, and a website has no compliance duty under it. What the chain does give you is a clear reason to know exactly what your own site currently says, because for the first time the main compliance route names the file you already have.

Start by reading it rather than assuming it. A robots.txt that was written for search engines in 2019 says nothing about the tokens that matter now, and a per-token evaluation is the only way to see the result rather than the intent: test robots.txt against each AI crawler and read the verdict for each one separately. The common surprise is not a rule that blocks too much. It is a file that was never updated and therefore allows everything, which is a decision nobody made.

Second, separate training from retrieval deliberately. Blocking a training crawler and blocking a retrieval crawler have different consequences, and only the second removes you from answers. If being cited matters, check what the platforms document about their retrieval fleets before you write a Disallow: the requirements for getting cited in ChatGPT and in Claude are published by their vendors and are not the same as the training tokens.

Third, check that allowing a crawler is enough. Access is only the first of two layers that decide whether AI can read your site, and a permissive robots.txt in front of a page that renders its text in the browser leaves a crawler with an empty document. Seeing what GPTBot actually receives settles that in one request, and it is the check most likely to change what you do next.

Finally, resist the urge to add a file because a regulation is in the news. The evidence on llms.txt is that publishing one changes nothing measurable about whether crawlers fetch it, and nothing in the AI Act, the Directive or the Code of Practice mentions it. The mechanisms those documents name are robots.txt, metadata and website terms. Two of those you already have, and the third is not a file you can add this afternoon.

  • Per-token robots.txt evaluation Which named AI crawler tokens are allowed and which are disallowed, evaluated per token rather than read as a whole file.
  • Training and retrieval separated Retrieval crawlers are what put you in an answer. Blocking them and blocking training tokens have different consequences.
  • Crawler view against browser view A permissive robots.txt in front of a client-rendered page still yields an empty document to a crawler.
  • EU rights reservation status Not a field any external scanner can fill. The duty sits on the model provider, and the evidence is inside their systems.
Four checks a site owner can actually run, each producing an observable result. Nothing here reports compliance with any EU instrument, because no external check can.

Related

Common questions

Does the EU AI Act require AI companies to respect robots.txt?

Not in those words. Regulation (EU) 2024/1689 contains the string robots.txt zero times. Article 53(1)(c) requires providers of general-purpose AI models to put in place a policy to identify and comply with a reservation of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790, naming no format. The file is named in the Commission's General-Purpose AI Code of Practice, published 10 July 2025, whose Measure 1.3 commits signatories to employ web crawlers that follow the Robot Exclusion Protocol as specified in IETF RFC 9309. Adhering to the Code is one route to demonstrating compliance, and the Code states that adherence is not conclusive evidence of it.

What happens on 2 August 2026 under the AI Act?

Article 113 states that the Regulation applies from 2 August 2026, with Chapter V applying from 2 August 2025 with the exception of Article 101. Article 53, the copyright policy obligation, is in Chapter V and has therefore applied since 2 August 2025. Article 101, which empowers the Commission to fine providers of general-purpose AI models up to 3 percent of annual total worldwide turnover or EUR 15 000 000, whichever is higher, falls to the general date. The obligation is not new on that date; the power to fine for breaching it is.

Does my website have to do anything to comply with the AI Act?

The obligations in Article 53 are on providers of general-purpose AI models, not on website owners, so a site has nothing to file and no status to declare. Article 4(3) of Directive (EU) 2019/790 makes a rightholder's reservation a condition on the text and data mining exception, which means expressing one is a choice a rightholder makes rather than a duty a site carries. Recital 18 of that Directive gives metadata and website terms and conditions as its examples of machine-readable means for online content.

Can a scanner tell me whether an AI company honoured my robots.txt?

No, and no external tool can. A scan observes the bytes your server returns, so it can fetch and evaluate your robots.txt for each crawler token and compare crawler and browser views of a page. Whether a particular model provider fetched that file and obeyed it happens inside their infrastructure and leaves no trace on your site. Server logs are weaker evidence than they look, because a user agent string is a claim that needs a reverse lookup against published address ranges to verify.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.