BlogFindings
The AI Act requires a machine-readable mark on AI text, and names no format
Article 50 of Regulation (EU) 2024/1689 has applied since 2 August 2026 and obliges providers of AI systems that generate synthetic text to mark the output in a machine-readable format. The Regulation names no format. Every technique it mentions sits in one sentence of one recital, inside a list introduced by the words such as, and the string HTML appears in the published text zero times.
Lantad has measured nothing here. This is a reading of a published legal text, quoted with links, reported rather than tested. What this site can add is the thing it exists to be precise about: the difference between an obligation placed on the system that wrote a page and a fact that fetching the page can establish. On this obligation those two do not overlap anywhere. AI visibility work starts from what a crawler receives when it asks for your URL, and nothing in Article 50 changes what a crawler receives.
In short
- Article 50(2) of Regulation (EU) 2024/1689 requires providers of AI systems generating synthetic audio, image, video or text content to ensure the outputs are, verbatim, marked in a machine-readable format and detectable as artificially generated or manipulated.
- The Regulation names no format for that mark. Searched over its published text on 21 August 2026, the words watermarks, metadata, cryptographic and fingerprints occur only inside one sentence of recital 133, a list opened by such as and closed by or other techniques, as may be appropriate.
- Article 113 states that the Regulation applies from 2 August 2026 and lists three exceptions, none of which covers Chapter IV, where Article 50 sits. Article 99(4)(g) makes a breach of Article 50 subject to administrative fines of up to EUR 15 000 000 or 3 percent of total worldwide annual turnover, whichever is higher.
- Article 50(4) requires a deployer publishing AI-generated text to inform the public on matters of public interest to disclose that fact, and exempts text that has undergone a process of human review or editorial control where a natural or legal person holds editorial responsibility for the publication.
- The string HTML appears zero times in the Regulation and every occurrence of the string http in it is part of an ELI citation URL in a footnote, so nothing in the text says where a mark on a published web page would live or how a crawler would read one.
Flow: AI system generates text (provider) to 50(2): mark in a machine-readable format; 50(2): mark in a machine-readable format to Deployer publishes the text; Deployer publishes the text (deployer) to 50(4): disclose to the public; Deployer publishes the text to AI crawler fetches the page; 50(4): disclose to the public to No format named for either; AI crawler fetches the page to No format named for either.
What Article 50 requires, and of whom
Article 50 sits in Chapter IV of Regulation (EU) 2024/1689, the chapter headed transparency obligations for providers and deployers of certain AI systems. It runs to seven paragraphs and they do four different jobs, on two different parties.
Paragraph 1 binds providers of AI systems intended to interact directly with natural persons. Those systems must be designed so that the person is informed they are interacting with an AI system, and the obligation lifts, verbatim, "unless this is obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect". That is the chatbot disclosure rule. It is about an interface a person is looking at, not about a document a machine is fetching.
Paragraph 2 is the one that reaches a published page, and it is short enough to quote whole. Verbatim: "Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated." The same paragraph asks that the technical solutions be "effective, interoperable, robust and reliable as far as this is technically feasible", and then carves out systems that perform "an assistive function for standard editing" or that "do not substantially alter the input data provided by the deployer or the semantics thereof". Text is named in that list alongside audio, image and video, so a model writing prose is inside the obligation and a spelling checker rewriting a sentence is outside it.
Paragraph 4 binds the deployer rather than the provider, and it is the one a publisher reads twice. Paragraphs 3, 5 and 7 handle biometric systems, the timing of the disclosure, and the codes of practice the AI Office is told to encourage. Paragraph 6 is a savings clause. The division that matters is between paragraph 2 and paragraph 4: paragraph 2 is an engineering obligation on the company that built the generator, and paragraph 4 is an editorial obligation on whoever publishes the result. Our earlier reading of the same Regulation on rights reservation followed a different chain through the law, from Article 53 to the Copyright Directive to the Commission's Code of Practice, and it arrived at a file on your server. This chain does not arrive on your server at all.
| Paragraph | Who it binds | What it requires | Does it reach a published page? |
|---|---|---|---|
| 50(1) | Provider | The person is informed they are interacting with an AI system | No, an interface rule |
| 50(2) | Provider | Output marked in a machine-readable format and detectable as artificially generated | Only if the mark survives publication |
| 50(3) | Deployer | Inform people exposed to emotion recognition or biometric categorisation | No |
| 50(4) | Deployer | Disclose deep fakes, and AI text published to inform the public on matters of public interest | Yes, as a disclosure to people |
| 50(5) | Both | Information given clearly and distinguishably at the latest at first interaction or exposure | Timing and clarity only |
| 50(7) | AI Office | Encourage codes of practice on detection and labelling of generated content | Not yet, no code is named in the text |
Where the format is named, and where it is not
The Article names no technique at all. Searched over the published text of the Regulation with markup removed, on 21 August 2026, the word watermarks appears once, metadata once and cryptographic once, and all three sit in a single sentence of recital 133. That sentence reads, verbatim: "Such techniques and methods should be sufficiently reliable, interoperable, effective and robust as far as this is technically feasible, taking into account available techniques or a combination of such techniques, such as watermarks, metadata identifications, cryptographic methods for proving provenance and authenticity of content, logging methods, fingerprints or other techniques, as may be appropriate."
Read that list for what it is. It opens with such as and closes with or other techniques, as may be appropriate, which makes it illustrative at both ends, and it lives in a recital. A recital states the reasoning behind the enacting provisions rather than the obligation itself, so the only place in this Regulation where a technique is named is the place where naming one binds nobody. The word fingerprints occurs twice in the whole text, once in that sentence and once in recital 30, where it means an actual human fingerprint in a passage about biometric categorisation. Two occurrences, two different senses of the word.
Nothing narrows it downstream either. The string HTML appears in the Regulation zero times. Every one of the six occurrences of the string http is part of an ELI citation URL in a footnote rather than a reference to the protocol. There is no header named, no media type named, no attribute named, and no cross-reference to any external specification that names one. The string robots.txt also appears zero times, which is the finding the earlier post on this Regulation was built around, and it is worth restating here only because it means neither of the Regulation's two machine-readable requirements is anchored to a concrete artefact anywhere in the enacted text.
That absence is defensible drafting rather than a lapse. A format written into primary legislation freezes at the moment of enactment and can only be changed by reopening the law, which is a poor fit for a marking technique that has to survive copying, quoting and reformatting. The consequence is still real for anyone downstream: until a code of practice under paragraph 7 or a harmonised standard fills the gap, two providers can both comply with paragraph 2 using marks that neither can read from the other.
The comparison worth drawing is with the opposite direction of travel. Where the question is a site telling machines what may be done with its content, there is a file everybody agrees on and an active standards effort refining what may be written in it: this blog has covered a preference vocabulary that cannot name a party and a draft that signs the preference into the content itself, both filed at the IETF, both revised this month. The AI preferences vocabulary draft is a working group document with a revision history anyone can read. Where the question is a machine telling readers and crawlers that it produced something, there is no equivalent file, no equivalent header, and no equivalent draft carrying the same standing.
One smaller detail, recorded because anyone repeating this count will hit it. Article 50(2) writes machine-readable with a hyphen and recital 133 writes machine readable without one. The hyphenated form occurs four times in the Regulation and the unhyphenated form twice, so a search for either alone returns a different answer, and nothing legal turns on the difference.
| String searched | Occurrences | Where they are |
|---|---|---|
| watermarks | 1 | Recital 133, inside the illustrative list |
| metadata | 1 | Recital 133, the same sentence |
| cryptographic | 1 | Recital 133, the same sentence |
| fingerprints | 2 | Recital 133, and recital 30 in the biometric sense |
| machine-readable | 4 | Article 50(2), plus the EU database and CE marking provisions |
| HTML | 0 | Nowhere in the text |
| robots.txt | 0 | Nowhere in the text |
When Article 50 started to apply, and what a breach costs
The date comes out of one article and three exceptions. Article 113 states, verbatim, that the Regulation "shall apply from 2 August 2026", and then lists what does not follow that date: Chapters I and II from 2 February 2025, Chapter III Section 4 together with Chapters V, VII and XII and Article 78 from 2 August 2025 with the exception of Article 101, and Article 6(1) with its corresponding obligations from 2 August 2027. Chapter IV appears in none of those three carve-outs, so Article 50 took effect on the general date, and it has been in application for nineteen days as this post is published.
The penalty sits in Article 99 and it is not the headline number. Article 99(3) reserves the largest fines, up to EUR 35 000 000 or 7 percent of total worldwide annual turnover, for the prohibited practices in Article 5. Article 99(4) sets the next band at up to EUR 15 000 000 or 3 percent of total worldwide annual turnover, whichever is higher, and lists seven provisions that fall into it. Point (g) of that list reads, verbatim: "transparency obligations for providers and deployers pursuant to Article 50". Article 99(6) then provides that for SMEs, including start-ups, each fine is capped at the lower of the percentage and the fixed amount rather than the higher.
Who imposes it is worth reading carefully, because the Regulation does not do it directly. Article 99(1) requires Member States to lay down the rules on penalties and other enforcement measures, which "may also include warnings and non-monetary measures", and to ensure they are effective, proportionate and dissuasive. So the ceiling is set in Union law and the machinery is national, which means the practical exposure of a publisher or an AI provider under Article 50 depends on the Member State as well as on the Regulation.
None of that is observable from outside an organisation, and this is the point at which a scanner has to be honest about its own reach. A fetch of your pages establishes what your server returned to a named user agent at a moment in time. It cannot establish which system wrote the text on those pages, whether that system marked its output, or whether any disclosure obligation was met. This blog has made the same distinction before about attestations of crawler behaviour, in the post on why a compliance record is falsifiable rather than proof, and the shape of the limit is identical here. What an external measurement can support, and what it cannot, is set out in full on our research page.
-
Application date2 August 2026 Article 113 sets the general date. Chapter IV, where Article 50 sits, is in none of the three exceptions. -
Fine ceilingEUR 15 000 000 or 3 percent Article 99(4)(g), whichever is higher. Article 5 prohibitions sit in a higher band at EUR 35 000 000 or 7 percent. -
SME treatmentLower of the two Article 99(6) caps each fine for SMEs and start-ups at whichever of the amount and the percentage is lower. -
Who enforcesMember States Article 99(1) requires Member States to lay down the rules on penalties, which may include warnings and non-monetary measures.
The carve-out for text that has been through editorial control
Paragraph 4 is where the Regulation touches anyone who runs a website with words on it, and it is narrower than the summaries of it usually suggest. The relevant sentence reads, verbatim: "Deployers of an AI system that generates or manipulates text which is published with the purpose of informing the public on matters of public interest shall disclose that the text has been artificially generated or manipulated."
Three conditions have to hold together before that bites. The text has to be generated or manipulated by an AI system, it has to be published, and its purpose has to be informing the public on matters of public interest. Product copy, documentation and marketing pages are not obviously any of the third, and the Regulation does not define matters of public interest anywhere in Article 50.
Then comes the exemption, and it is the sentence to read closely: the obligation "shall not apply where the use is authorised by law to detect, prevent, investigate or prosecute criminal offences or where the AI-generated content has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility for the publication of the content". The second limb has two conditions joined by and. Human review or editorial control is not enough on its own; somebody identifiable has to hold editorial responsibility for the publication. A site that runs generated drafts past a named editor who owns what ships is inside the exemption. A site that publishes generated text with nobody holding that responsibility is not, and the reviewing step alone will not rescue it.
What paragraph 4 does not say is as important as what it does. It does not require the disclosure to be machine-readable. The machine-readable requirement in Article 50 is paragraph 2 and it binds the provider of the generating system. Paragraph 4 binds the deployer and asks for a disclosure, and paragraph 5 says that the information in paragraphs 1 to 4 "shall be provided to the natural persons concerned in a clear and distinguishable manner at the latest at the time of the first interaction or exposure" and "shall conform to the applicable accessibility requirements". Natural persons. The audience for the paragraph 4 disclosure is a reader, not a crawler, and no part of Article 50 asks a publisher to emit anything a parser could pick up.
That gap matters for anyone trying to work out how much of the web an answer engine is now reading back to itself. The audit this blog reported on AI-generated sources inside AI search citations had to reach for a detection classifier precisely because no label existed to count, and its own authors treated their figure as a lower bound for that reason. Article 50 does not close that gap for text on a page: it puts a mark inside the provider's output pipeline and a sentence in front of a human reader, and leaves the space between them empty. The same shape appeared when this blog looked at whether review authenticity is visible in the markup, where the thing a reader wants to verify is exactly the thing the markup does not carry.
Search engine documentation does not fill the space either, and it does not claim to. Google's page on how content appears in AI features, carrying Last updated 2025-12-10 UTC, states that the best practices for SEO remain relevant for AI features and that there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary". No crawler vendor documentation read for this blog asks for a generation label, credits one, or penalises its absence. So a disclosure under paragraph 4 is a legal obligation and, on the published evidence available today, not a generative engine optimisation input in either direction.
Disclosure required
- Text generated or manipulated by an AI system
- The text is published
- Published with the purpose of informing the public
- On matters of public interest, a term Article 50 does not define
- Disclosure is to natural persons, clear and distinguishable
Exempt
- Human review or editorial control has taken place
- AND a natural or legal person holds editorial responsibility
- Or the use is authorised by law for criminal enforcement
- Artistic, creative, satirical or fictional works, limited disclosure
- Nothing here requires a machine-readable label
What an external scan of your site can see, which is none of this
This is the part where a scanner has to say what it does not do, so here is the inventory from the code rather than from a claim. The StructureChecks interface in core/src/verdict.ts holds exactly seven booleans: a title, a meta description, exactly one h1, a canonical, heading coverage, no oversized block, and llms.txt present. None of them concerns provenance, authorship, or how the text was produced. The composite score has four components and their weights are settings rather than findings: parity at 0.50, access at 0.25, structure at 0.15 and schema at 0.10, all four declared in core/src/config.ts. There is no fifth component and no place to add a generation label to the existing four without changing what the grade means.
The extractor is narrower still. Reading core/src/extract.ts on 21 August 2026, it collects six HTML attributes across the whole document: type, lang, name, content, rel and href. A provenance attribute would not be read even if one existed and even if a page carried it, which means this scanner could not report an Article 50 mark today under any format the Union might eventually settle on. That is a limit worth stating plainly rather than a roadmap item, because the format does not exist to build against.
What a scan can establish is the layer underneath all of it, and that layer is unaffected by any of this. Whether a named crawler is allowed by your robots.txt, which you can resolve per token with the robots.txt tester. What a specific user agent actually receives when it asks, which is what the GPTBot view reports. Which tokens exist to be addressed at all, listed on our AI crawlers reference. Whether the text a human sees survives into what a fetch returns, which is the prose parity measure. Whether the structured data on the page parses. How those pieces combine into a grade is set out on the methodology page, and the conduct of our own fetcher is published at our bot policy.
Set the two machine-readable requirements side by side and the asymmetry is the finding. For a site expressing a preference about its content there is a protocol, and RFC 9309 is candid in its own text that the rules it standardises are not a form of access authorization, which is a limit you can at least read and plan around. For an AI system marking its own output there is no protocol at all, only an obligation and a recital full of examples. A site owner can check the first from outside in a few seconds. Nobody, including a regulator, can check the second by fetching a URL.
The practical advice that follows is short, and it is smaller than the subject. If you publish text on matters of public interest and any of it is generated, decide who holds editorial responsibility for it and write that down, because that is the condition the exemption turns on and it is an organisational fact rather than a technical one. Do not go looking for a tag to add, because there is not one to add and inventing your own would signal nothing to anybody. And keep the layer you can actually control in good order, because whether an AI crawler can obtain and read your page still decides whether any of this is ever reached, which is the same reasoning behind everything on getting cited in Google AI Overviews.
- Whether a named crawler is allowed by robots.txt Resolved per token against the file your server returns, which is the layer an external fetch can establish.
- What a specific user agent receives from your origin The status, the body and whether the text a human sees survives into it. This is the parity component, weighted 0.50.
- Whether the page's structured data parses The schema component, weighted 0.10 in core/src/config.ts. A setting somebody chose, not a measured importance.
- Whether any text on the page was generated by an AI system Not read, not inferable from a fetch, and not carried by any attribute the extractor collects.
- Whether an Article 50(2) mark is present No format exists to look for. The extractor reads six attributes: type, lang, name, content, rel and href.
- Whether an Article 50(4) disclosure was made The disclosure is addressed to natural persons and the Regulation asks for no machine-readable form of it.
Lantad
Published .
A rule that tells a machine to look for something has to say what the thing looks like. Chapter IV of the EU AI Act now obliges the providers of systems that generate text to mark what those systems produce, in a format a machine can read. The Regulation says that in one sentence and then stops. It names no header, no file, no attribute and no syntax, and the only place any technique appears at all is a recital, in a list of examples that is open at both ends.
Common questions
Does the EU AI Act require me to label AI-generated text on my website?
Only in a narrow case. Article 50(4) applies to a deployer publishing AI-generated or AI-manipulated text with the purpose of informing the public on matters of public interest, and it exempts content that has undergone human review or editorial control where a natural or legal person holds editorial responsibility for the publication. Both limbs of that exemption have to hold. The separate machine-readable marking duty in Article 50(2) binds the provider of the AI system, not you.
What format does the AI Act require for a machine-readable mark?
None is specified. Article 50(2) requires the output to be marked in a machine-readable format and detectable as artificially generated or manipulated, and says the solution should be effective, interoperable, robust and reliable as far as technically feasible. Recital 133 names watermarks, metadata identifications, cryptographic methods, logging methods and fingerprints as examples, in a list opened by such as and closed by or other techniques, as may be appropriate. Article 50(7) leaves the detail to codes of practice the AI Office is to encourage.
When did Article 50 start to apply?
2 August 2026. Article 113 states that the Regulation applies from that date and lists three exceptions covering Chapters I and II, Chapter III Section 4 with Chapters V, VII and XII and Article 78, and Article 6(1). Chapter IV, which contains Article 50, is in none of them, so it took effect on the general date.
Can Lantad tell me whether a page carries an AI generation mark?
No, and it could not even if a format existed. The extractor in core/src/extract.ts reads six HTML attributes, being type, lang, name, content and rel and href, and the seven structure checks in core/src/verdict.ts cover a title, a meta description, one h1, a canonical, heading coverage, oversized blocks and llms.txt. Nothing in the scan or the score concerns provenance. What a scan does establish is whether a named AI crawler can reach and read the page at all.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.