# The AI preferences standard you cannot deploy yet

> The IETF vocabulary for expressing AI usage preferences reached its sixth draft on 28 April 2026. The companion draft that defines where to write one down expired on 1 May 2026, so there is still nothing on your site for a crawler to read.

- Canonical page: https://lantad.co/blog/the-ai-preferences-standard-you-cannot-deploy-yet
- This file: https://lantad.co/blog/the-ai-preferences-standard-you-cannot-deploy-yet.md
- Last substantive update: 2026-07-28

## Key facts

- **Published:** 2026-07-28
- **Category:** Guides
- **Author:** Lantad
- **Length:** 4252 words
- **Takeaway 1:** The IETF AI Preferences working group's vocabulary draft, draft-ietf-aipref-vocab-06, was published on 28 April 2026 on the Standards Track and defines exactly two usage categories: AI model training and search.
- **Takeaway 2:** The companion draft that defined where those preferences are written, draft-ietf-aipref-attach-04, was published on 28 October 2025 and carries an expiry of 1 May 2026, so the Content-Usage HTTP header and the robots.txt content-usage rule it specified currently sit behind an expired document.
- **Takeaway 3:** The vocabulary draft states in its own text that enforcement is not provided by the specification and that preferences do not themselves create rights, obligations, or prohibitions.
- **Takeaway 4:** Until an attachment mechanism is published, robots.txt as described by RFC 9309 and the per-crawler product tokens each vendor documents remain the only preference signals an AI crawler reads on a site today.
- **Takeaway 5:** Lantad has not measured AIPREF adoption and reports the state of these documents from the IETF datatracker as read on 28 July 2026, not from any scan.

## Summary

Every few months somebody announces that there is now an AI preferences standard, that it settles the argument about what AI systems may do with your content, and that the thing to do next is implement it. The IETF's AI Preferences working group is the serious version of that story. It was chartered in February 2025, its vocabulary is on the Standards Track, and the people writing it come from Mozilla, Open Future and Microsoft rather than from a marketing department. It is the closest thing to a real answer that exists.

What it does not have, as of 28 July 2026, is a published way to write a preference anywhere [an AI crawler](https://lantad.co/glossary/ai-crawler) would look for it. The vocabulary that says what you can express is progressing. The document that said where to put it lapsed almost three months ago. This post reports the state of [the working group's documents](https://datatracker.ietf.org/wg/aipref/documents/) as listed on the IETF datatracker, then covers the part that is measurable from outside a site, which is what a crawler actually reads today while the standard settles. Lantad has not measured AIPREF adoption, and at present there is nothing to measure: a vocabulary with no attachment mechanism leaves no trace on a website that a scanner could find.

## What the IETF AI preferences working group is actually building

The working group was announced on the IETF's own blog on 27 February 2025, at https://www.ietf.org/blog/aipref-wg/, and the reasoning in that post is worth repeating because it is the reason any of this exists. Vendors had settled on what the post calls a confusing array of non-standard signals in the robots.txt file and elsewhere to guide their crawling and training decisions. Publishers had lost confidence that stated preferences were being honoured, and were falling back on blocking IP addresses instead. The group was chartered to replace that with two things: a common vocabulary for expressing preferences about use of content for AI training and related tasks, and a standard way of attaching that vocabulary to content.

Two building blocks, and they are genuinely separable. A vocabulary without an attachment mechanism is a dictionary for a language nobody can write down. An attachment mechanism without a vocabulary is an envelope with nothing to put in it. The working group split the work along exactly that line and produced one draft for each.

Read the datatracker today and the split is visible in the states. [The vocabulary draft](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/) is at revision 06, dated 28 April 2026, thirteen pages, active, and marked Proposed Standard. [The attachment draft](https://datatracker.ietf.org/doc/draft-ietf-aipref-attach/) is at revision 04, dated 28 October 2025, ten pages, and marked expired. One half of the charter has a live document behind it and the other half does not.

Around those two sit a set of individual submissions that have not been adopted as working group documents. [Crawler best practices](https://datatracker.ietf.org/doc/draft-illyes-aipref-cbcp/), dated 9 April 2026, is written by Gary Illyes, Mirja Kühlewind of Ericsson and AJ Kohn, and covers what automated HTTP clients should do when they access web resources, including documenting crawler identity and data usage transparently. [A competing vocabulary from Microsoft](https://datatracker.ietf.org/doc/draft-madhavan-aipref-displaybasedpref/), dated 25 March 2026 and authored by Krishna Madhavan, Fabrice Canel, J. Gimbel and S. Cooper, proposes terminology for controlling how search engines and AI systems use collected content, and is listed as a candidate for working group adoption rather than an adopted document. There is also a verifiable credential binding dated 2 July 2026 and an exclusions draft dated 22 April 2026.

The distinction between an adopted working group document and an individual submission matters more than it looks, and it is the first place a summary of this work usually goes wrong. Every Internet-Draft carries boilerplate stating that it is not endorsed by the IETF and holds no formal standing in the standards process. For an adopted document that boilerplate is a formality on the way to something. For an individual submission it is a plain description of the situation: one or more people have written down a proposal and asked a room to look at it. Anyone selling you compliance with a draft in the second category is selling you compliance with a document. If you are trying to understand [generative engine optimisation](https://lantad.co/glossary/geo) from primary sources, reading the state field next to each document name is most of the work, and [the working group's document list](https://datatracker.ietf.org/wg/aipref/documents/) prints it next to every title.

## The vocabulary defines two categories, and that is the whole vocabulary

Open [draft-ietf-aipref-vocab-06](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/) expecting a taxonomy and the first surprise is how small it is. Thirteen pages, and the entire vocabulary is two usage categories.

The first is AI model training, defined in the document as the act of using an asset in the production or refinement of an AI model that can generate content in one or more modalities, with text, image and audio given as the examples. The second is search, defined as use of an asset in an application where the primary purpose of the application is to select assets and direct users to the location of those assets. Each serialises to a short label with a binary value: train-ai set to y or n, and search set to y or n. That is the complete surface.

Two things follow from that shape, and the second is the one that gets missed. The first is that the split maps onto the distinction the crawler vendors already draw in their own fleets. [OpenAI's crawler documentation](https://developers.openai.com/api/docs/bots) lists four tokens with four separate purposes: OAI-SearchBot to surface websites in search results in ChatGPT's search features, GPTBot to make its generative foundation models more useful and safe, ChatGPT-User for certain user actions in ChatGPT and Custom GPTs, and OAI-AdsBot to validate the safety of pages submitted as ads. [Google's crawler documentation](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) draws a comparable line between crawling for Search and the separate product token that governs generative training. Training against search is the division the market has already converged on, so standardising exactly that division is a reasonable place to start rather than a failure of ambition.

The second is that the standard is deliberately smaller than the vocabulary the market is using, and a two value split cannot express most of what site owners currently argue about. Cloudflare's July 2026 announcement sorted bot traffic into eleven classifications including Agent, Transact, Data Collection, SEO and Ads Verification, which we covered when [the edge and the file disagreed](https://lantad.co/blog/two-layers-decide-if-ai-can-read-your-site). An agent fetching a page on behalf of a user who asked a question is not obviously training and not obviously search. Retrieval augmented generation, where a model fetches a page at answer time and quotes it, sits awkwardly between the two definitions. None of that is an oversight: a standard that tried to name every commercial arrangement in this market would be obsolete before it shipped, and the group has been trimming rather than adding. The [minutes of its interim meeting on 14 April 2026](https://datatracker.ietf.org/doc/minutes-interim-2026-aipref-01-202604141315/) record agreement to remove the word large from the text, on the grounds that using it would require a quantified threshold for model size that nobody thought workable, along with agreement to drop the high level definitions of artificial intelligence and machine learning as not technically required.

The practical consequence for anyone working on [AI visibility](https://lantad.co/glossary/ai-visibility) is that this vocabulary, even fully deployed, would not answer the question most people want answered. It expresses whether you permit training and whether you permit search. It does not express whether you want to be cited, on what terms, or with what attribution, which is the actual commercial question underneath [answer engine optimisation](https://lantad.co/glossary/aeo). The document is honest about its own scope in a way the coverage of it usually is not.

## Why there is nowhere to write an AI preference down yet

The expired half is the one that would have touched your server, and it is worth knowing what it said because it is the thing people believe already exists.

[draft-ietf-aipref-attach-04](https://datatracker.ietf.org/doc/draft-ietf-aipref-attach/), titled Associating AI Usage Preferences with Content in HTTP, defined two attachment points. One was an HTTP response header field named Content-Usage, registered as a structured field dictionary, so a server could answer a request with Content-Usage: train-ai=n alongside the content itself. The other was a new rule inside robots.txt, spelled content-usage, taking an optional path pattern followed by a usage preference, so a file could carry a line reading content-usage: /ai-ok/ train-ai=y and scope the preference to part of a site. The draft's header block states that it updates [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html), the Robots Exclusion Protocol, if approved, which is what makes the robots.txt half more than a convention: it would have been a sanctioned extension to the file crawlers already parse.

That document carries a published date of 28 October 2025 and an expiry of 1 May 2026, roughly six months apart, and the datatracker has listed it as expired since. It is important to be precise about what that does and does not mean, because the word invites a conclusion it does not support. An expired Internet-Draft has not been rejected, withdrawn or voted down. Expiry is a calendar fact: a draft that is not replaced by a newer revision within its lifetime stops being current, and the working group can revive it with a new revision at any time, which is an ordinary event rather than a rescue. Plenty of documents that became RFCs expired at least once on the way.

What it does mean, precisely and only, is this: on 28 July 2026 there is no live specification defining where an AIPREF preference goes. The vocabulary is a set of terms with no sanctioned home. You can serve a Content-Usage header today, because a server may send any header it likes, and it will cost you nothing and do nothing, because no crawler is obliged to parse a header whose defining document is not current and none has published that it does.

This is a shape worth recognising, because the site has already published on it once. [The evidence on llms.txt](https://lantad.co/blog/what-the-evidence-says-about-llms-txt) reports an Ahrefs study of 137,210 domains measured in May 2026 in which 97 percent of valid llms.txt files were never fetched at all. That is a proposal in wide deployment with no measured effect. AIPREF's attachment mechanism is the opposite failure, a well designed mechanism with no deployment, and both produce the same outcome for a site owner: a file or a header that changes nothing about what a crawler does. The difference between [llms.txt](https://lantad.co/glossary/llms-txt) and Content-Usage is not quality of design. It is that neither is being read, and only one of them has a live document behind it. The only preference file with a specification a crawler is documented to parse remains robots.txt, which you can check against a specific token with [the robots.txt tester](https://lantad.co/tools/robots-txt-tester).

## A preference is not an access control, and the draft says so in its own text

The most useful paragraph in the vocabulary draft is not about vocabulary. It is the passage stating that enforcement is not provided by this specification, that preferences do not themselves create rights, obligations, or prohibitions, and that other mechanisms, technical, legal, contractual or otherwise, might enforce adherence or non-adherence to these preferences and thereby determine the consequences of not respecting a stated preference.

That is a specification declining to promise what its readers most want it to promise, written into the document rather than left to a FAQ. It puts AIPREF in the same category as the file it would extend. [The Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html), published on the Standards Track in September 2022, is explicit that its rules are not a form of access authorization, which is the point [the two layers post](https://lantad.co/blog/two-layers-decide-if-ai-can-read-your-site) was built on: a file that permits a crawler cannot guarantee that crawler is served, and a file that forbids one cannot guarantee it is refused.

There is a second gap in the same area, and it is the one that makes the first one harder. A preference is expressed to whoever is asking, and the identity of the asker is itself a claim. A request that arrives announcing itself as GPTBot has asserted a user agent string and nothing more, which is why [a user agent is a claim rather than an identity](https://lantad.co/blog/a-user-agent-is-a-claim-not-an-identity) and why verifying one costs a reverse DNS or published IP range check that most tooling never performs. Any preference vocabulary keyed to a crawler's declared identity inherits that problem whole. A rule saying train-ai=n for one named agent is a rule addressed to whoever chooses to say they are that agent.

None of this makes the work pointless, and it would be easy to write a cynical version of this post that concluded standards do not matter. That conclusion does not follow. A shared vocabulary has value precisely because the alternative, described in the working group's own charter announcement, is every vendor inventing its own signal and publishers giving up and blocking IP ranges. Standardising the words is a real contribution even when nothing enforces them, and the difference between one agreed term and nine vendor specific ones is the difference between a preference a crawler could honour and a preference it would have to go looking for.

What it does mean is that the honest description of AIPREF, today, is a request rather than a control, and that the layer which actually decides whether a crawler is served is the response your infrastructure returns. That distinction is not a detail for a site owner. It is the difference between believing a page is protected and knowing what happened to a request, and it is why [the way Lantad scores a page](https://lantad.co/methodology) grades the response that came back rather than the intent of the file that was published. Cloudflare's [announcement of default blocking for Training and Agent bots](https://blog.cloudflare.com/content-independence-day-ai-options/) will move more traffic than any preference vocabulary will, because it operates on the layer that answers.

## What an AI crawler actually reads on your site today

Strip out everything that is proposed, drafted or expired, and the list of preference signals an AI crawler is documented to read on your site right now is short enough to hold in your head.

There is robots.txt, parsed under RFC 9309, containing user agent groups and allow and disallow rules. There are the product tokens each vendor documents for its own fleet, which is where the per purpose control actually lives today, and which is why the vendor documentation pages are the primary sources rather than any standard. And that is the list. Everything else is either a proposal, a header nothing parses, or a control that lives at a layer robots.txt cannot see.

Within that short list there are traps that have nothing to do with AIPREF. Some of the tokens you are told to add are robots.txt product tokens that send no user agent of their own, so a request from them never appears in a log under that name and searching for one returns zero on every site, forever, which we covered in the post on [tokens that never appear in your logs](https://lantad.co/blog/the-crawler-tokens-that-never-appear-in-your-logs). A zero there is not evidence of anything. Tokens are also not interchangeable between vendors, so a rule written for one fleet does nothing for another, and the per platform notes on getting cited by [ChatGPT](https://lantad.co/how-to-get-cited/chatgpt) and by [Claude](https://lantad.co/how-to-get-cited/claude) exist because the answer differs by platform rather than because the guidance is duplicated.

The useful way to think about the gap is that AIPREF, if and when it lands, would standardise the words in the file. It would not change the two facts underneath: that a preference is honoured voluntarily, and that a crawler which is permitted still has to be able to read the page once it arrives. The second is the one this product exists to measure, and it is unaffected by anything the working group publishes. A site that returns an empty shell to a crawler because its content renders client side is invisible to that crawler whether its robots.txt says train-ai=y, train-ai=n, or nothing at all. The permission and the readability are separate questions and a standard settles only the first.

If you want to see which of the documented crawlers your own site currently admits, [the AI crawler check](https://lantad.co/tools/ai-crawlers) tests robots.txt per token rather than as one verdict, which matters because the common failure is a file that admits one fleet and refuses another without anyone intending it. For the other side of the exchange, [LantadBot's conduct policy](https://lantad.co/bot) documents the token this scanner sends and how to refuse it, on the principle that a tool asking you to audit crawler behaviour should be auditable itself.

## What to check on your own site while the standard settles

The reason to follow standards work is to know what not to do yet. Here is what that means in practice, and none of it is waiting.

Do not deploy a Content-Usage header and count it as protection. Sending it is harmless and costs nothing, and if the draft is revived you will already be compliant, so there is a defensible argument for adding it. What you must not do is record it as a control in a risk register, tell a client their content is protected from training, or let it substitute for a robots.txt rule. A header that nothing parses is a note to yourself. The same caution applies to any vendor currently selling AIPREF compliance: ask which document they are compliant with and check its state on the datatracker, which takes about thirty seconds and settles the question.

Do check that your robots.txt parses and that you know what each token resolves to. This is unglamorous and it is where the real failures are: a rule written for a token that does not exist, a group that shadows another because of ordering, an allow that never applies. Testing a specific token against your live file is the check that catches all three, and [testing per token rather than per file](https://lantad.co/tools/robots-txt-tester) is the only way to see it, because a file that looks correct as a whole can still refuse one fleet.

Do check the layer below the file. Whatever robots.txt says, the thing that decides the outcome is the response an unauthenticated request receives, and that is set by your CDN, your firewall and your bot management rules. Since Cloudflare's new defaults apply from 15 September 2026 to newly onboarded domains, a site moved onto a new zone between now and then can acquire a blocking posture nobody at the company chose and nothing in the repository records.

Do check that a crawler you have admitted can actually read the page. This is the part that no standard touches and the part that is most often broken. If the crawler gets the navigation and the footer and none of the article, permission was never the constraint. That gap between what a browser renders and what a crawler receives is [prose parity](https://lantad.co/glossary/prose-parity), and it is the single most common reason a page that is fully permitted is still not readable. It is a per stack problem more than a per site one, which is why the remedies for [a Next.js site](https://lantad.co/fix/nextjs) differ from those for a client rendered app, and it is measured on the response rather than inferred from the framework.

Finally, be careful about what you conclude from a check that did not complete. If a fetch times out or is refused, the honest report is that the measurement failed, not that the site is blocking AI, and the difference between those two statements is why [we withhold a grade rather than print a guess](https://lantad.co/blog/why-we-withhold-a-grade) when a scan degrades. That principle applies to reading standards work as much as to reading a scan. The state of AIPREF on 28 July 2026 is that one document is active, one has expired, and several are proposals nobody has adopted. That is a smaller claim than the coverage of it usually makes, and it is the one the datatracker supports.

## Questions and answers

**Is there an AI preferences standard I can implement today?**

No. The IETF vocabulary draft, draft-ietf-aipref-vocab-06, was published on 28 April 2026 and is active, but the companion draft defining where to write a preference down, draft-ietf-aipref-attach-04, expired on 1 May 2026. Until an attachment mechanism is published there is no sanctioned place to put an AIPREF preference, so robots.txt and the per-crawler product tokens each vendor documents remain the only signals a crawler reads.

**What does draft-ietf-aipref-vocab actually define?**

Two usage categories and nothing else. AI model training, described as using an asset in the production or refinement of an AI model that can generate content, and search, described as use of an asset in an application whose primary purpose is to select assets and direct users to them. Each serialises to a binary value, train-ai=y or n and search=y or n. In the absence of a statement of preference the draft assigns every category the value unknown, and it takes no position on what default might be applied.

**Does an expired Internet-Draft mean the work was rejected?**

No. Expiry is a calendar event, not a verdict. A draft that is not replaced by a newer revision within its lifetime stops being current, and a working group can revive it with a new revision at any time. draft-ietf-aipref-attach-04 carries a published date of 28 October 2025 and an expiry of 1 May 2026. What its expiry does mean is that there is no live specification defining where an AIPREF preference goes.

**Should I add a Content-Usage header to my site now?**

You can, and it will do nothing. A server may send any header, so sending Content-Usage costs nothing and would put you ahead if the attachment draft is revived. What it must not do is count as a control: no crawler has published that it parses the header, its defining draft expired on 1 May 2026, and recording it as protection in a risk register or telling a client their content is protected would be a claim the evidence does not support.

---

Lantad measures whether AI crawlers can actually read a page: it fetches as a non-rendering
crawler, renders as a browser, and reports the gap. Free scan, one URL, no signup.

Method and weights: https://lantad.co/methodology | All pages as markdown: https://lantad.co/md | Crawler policy: https://lantad.co/bot
