BlogFindings

An AI usage preference cannot name a party, and two drafts propose fixes

The IETF vocabulary for AI usage preferences defines two categories and two values, y and n, and offers no way to scope either to a named company. Two individual submissions filed on 6 August 2026 propose two different places to put the missing permission.

16 min read Lantad

The IETF working group building a standard vocabulary for this has scoped itself deliberately, and an earlier post here covered the piece you cannot deploy yet, which is the companion draft defining where a preference gets written down. This post is about a different limit in the same specification, and shipping the missing attachment mechanism would not close it: the vocabulary has no concept of a party at all. On 6 August 2026 two individual submissions arrived at the IETF proposing two different answers to that gap, one substantial and one four pages long. Lantad has measured nothing here and holds no data on how any of it behaves in the wild, because none of it is deployed anywhere. Everything below is read from the drafts and the RFCs themselves on 15 August 2026, so nothing in this post is a change to make today to what an AI crawler sees on your site.

In short

  • The AI preferences vocabulary at draft-ietf-aipref-vocab-06, dated 28 April 2026, defines exactly two usage categories, AI Model Training and Search, carrying the labels train-ai and search, and exactly two preference values, y and n.
  • That vocabulary has no field for a party, so a preference is broadcast to every recipient at once, and its conflict rule states that where statements disagree the most restrictive preference applies.
  • draft-wallace-aipref-grant-binding-01, filed on 6 August 2026, proposes a signed credential that lifts a preference for one named grantee over one asset for a stated period, and states that a grant must identify the asset by content digest.
  • draft-zehta-aipref-parameters-00, filed the same day by Creative Commons, proposes attaching structured field parameters to a preference, with a worked example of a tip jar path carried alongside a train-ai value.
  • Lantad measured none of this. Every figure and quotation here is read from the two drafts, the vocabulary draft and RFC 9309 on 15 August 2026, and no mechanism described in this post is deployable on any site today.
DocumentDateStatusWhat it expressesNames a party
draft-ietf-aipref-vocab-0628 April 2026Working group, standards trackAllow or disallow, per usage categoryNo
draft-wallace-aipref-grant-binding-016 August 2026Individual, standards trackA signed grant over one asset for one granteeYes
draft-zehta-aipref-parameters-006 August 2026Individual, informationalParameters attached to a preference valueNo
RFC 9309September 2022Published standards trackAllow and disallow paths, per product tokenToken only
The three documents this post reads, with the dates they carry, read at the IETF datatracker on 15 August 2026. The party column records whether the mechanism can scope a statement to one named recipient.

What an AI usage preference can actually express

The vocabulary lives in draft-ietf-aipref-vocab-06, dated 28 April 2026, edited by Paul Keller of Open Future and Martin Thomson of Mozilla. Its abstract describes a vocabulary for expressing preferences about how digital assets are used by automated processing systems, allowing the declaration of restrictions or permissions for such use. Two things in that sentence are worth holding on to. The word permissions is there, so the design is not refusal only. And the subject of the sentence is the asset, never the recipient.

The vocabulary defines two categories of use and no more. AI Model Training is defined as the act of using an asset in the production or refinement of an AI model that can generate content in one or more modalities. Search is defined as use of an asset in an application whose primary purpose is to select assets and direct users to their location, and the definition carries conditions: the presentation must include a direct reference or link back to where the asset came from, and displayed excerpts must serve to help users evaluate relevance. The draft then adds a sentence that decides a great deal, stating that this category does not include the use of assets to generate summaries. An AI Overview is not Search under this vocabulary.

The serialization is deliberately small. Section 6 maps the categories to the labels train-ai and search, maps the two preferences to the single byte tokens y and n, and gives the whole worked example in one line: train-ai=y, search=n. After processing, a recipient assigns each category one of three values, allowed, disallowed, or unknown, and in the absence of any statement every category is unknown.

That is the entire expressive range, and it is a narrower thing than the files site owners are used to. It is narrower than the RSL licence directive, which writes a licence into robots.txt and changes no access. It is narrower than the W3C TDM reservation protocol, which no major crawler documentation names. It is also, in one specific respect, narrower than the robots.txt file that already sits on every site, which is the subject of the third section below.

CategoryLabelAllowed valuesResolved states
AI Model Trainingtrain-aiy or nallowed, disallowed, unknown
Searchsearchy or nallowed, disallowed, unknown
Any named companyNot expressibleNot expressibleNot expressible
The complete vocabulary as defined in sections 4 and 6 of draft-ietf-aipref-vocab-06, read on 15 August 2026. Two categories, two labels, two values. There is no third column for a recipient because the data model has no field for one.

Why a preference resolves to the strictest answer available

Because a preference has no addressee, more than one of them can apply to the same asset, and the specification has to say what happens then. Section 5.1 of the vocabulary draft handles it in three bullets. If any statement of preference says the usage is disallowed, the result is disallowed. Otherwise, if any statement allows it, the result is allowed. Otherwise the preference is unknown. The draft states the intent plainly in the following sentence: this process ensures that the most restrictive preference applies.

That rule is correct for the problem it solves and it forecloses the problem this post is about. A negotiated permission is by construction a less restrictive statement than the standing reservation it sits beside, so under a strictest-wins rule it can never win. The specification is aware of this, and the escape hatch it names is not technical. Section 5.2 says that contractual agreements or other specific arrangements might override statements of preference, and expands: where arrangements such as legal agreements explicitly permit the use of an asset, those arrangements likely apply despite the existence of machine-readable statements of preference. The permission exists. It is simply somewhere no machine can read it.

The draft is candid about its own standing in a way that is worth quoting rather than paraphrasing. It carries a Note to Readers stating that its contents do not reflect consensus of the working group either in whole or part, and sections 3 and 4, which are the data model and the vocabulary itself, each open with a note saying the section does not yet have consensus. Neither the category list nor the two-value model is settled.

It is candid about enforcement too. The draft states that enforcement is not provided by the specification, and that preferences do not themselves create rights, obligations, or prohibitions. This is the same division of labour that runs through everything in this area: the file states a position, and something else entirely decides what happens. It is why the EU AI Act never writes your robots.txt for you, and why two separate layers decide whether an AI system can read your site.

Resolving a usage category, following sections 5 and 5.1 of draft-ietf-aipref-vocab-06. A description of the specified algorithm, not a measurement of any implementation.

Robots.txt names a product token, which is not a company

The obvious objection is that robots.txt already solves this, because a user-agent group addresses a specific crawler. It does, and the distinction between what it addresses and what a grant would address is the cleanest way to see the gap.

RFC 9309, the Robots Exclusion Protocol, published in September 2022 by Martijn Koster, Gary Illyes, Henner Zeller and Lizzi Sassman, defines the user-agent line in section 2.2.1 with a sentence that gives the whole game away: crawlers set their own name, which is called a product token, to find relevant groups. The name in your file is a name the crawler chose for itself and can change. It is not an identity, it is not authenticated, and the specification requires nothing of it beyond a character set and a recommendation that it appear as a substring of the User-Agent header.

Everything that follows from that has been visible in practice for a while. A user agent is a claim rather than an identity, which is why verification schemes exist at all. A renamed token leaves a robots.txt group matching nothing, silently, with no error anywhere. One company commonly ships several tokens, and six of the nine vendors in Lantad's registry publish exactly one, so the mapping between a token and a corporate counterparty is neither complete nor stable in either direction.

Then there is what the rules mean. RFC 9309 disposes of that in one line at the end of its introduction: these rules are not a form of access authorization. Its security considerations repeat the point, saying the protocol is not a substitute for valid content security measures. An Allow line is a request that a crawler is asked to honour, not a permission granted to a party, and the difference is exactly the difference between the two documents filed on 6 August. This is also why a robots.txt reading tells you less than it appears to, and why a file that blocks every citation crawler can still grade B on a score that weights four things and not one.

robots.txt user-agent line

  • Addresses a product token
  • The crawler sets its own name
  • Unauthenticated, changeable
  • Rules are not access authorization
  • Public, readable by anyone
  • No expiry, no withdrawal record

A grant credential

  • Addresses a named grantee
  • The rights holder signs it
  • Verified against an issuer key
  • Lifts a preference for that grantee
  • Held by the grantee, not published
  • Validity period plus revocation
What a robots.txt user-agent line addresses, as specified in RFC 9309 section 2.2.1, against what the grant draft addresses in its section 2. Both descriptions are read from the documents, not measured.

What the grant draft would add, and what it says it cannot

draft-wallace-aipref-grant-binding-01, filed on 6 August 2026 by an author at LicenseFoundry, states the gap in its own abstract better than a summary can. A preference expresses a reservation, and it does not, by itself, provide a verifiable, revocable record of a specific grant that lifts a preference for a specific party. The draft calls itself a discussion starter and says the mechanism is one candidate, not a final design.

The mechanism is a Verifiable Credential secured as a JSON Web Signature, scoped to a tuple of asset, category, grantee and validity period. It carries a resolvable identifier for the grant, an issuer identifier a verifier can resolve to a public key, the asset identified by a cryptographic content digest, the named grantee, one or more usage entries pairing an AIPREF category with a granted value, a validity window expressed as validFrom and validUntil, and a revocation reference such as a status list entry. The draft's requirements are short: a grant must be signed by the grantor's key, must identify the asset by content digest, and should carry a revocation reference, because a grant without one cannot be withdrawn after issuance.

Verification is four steps and needs no contact with the grantor beyond fetching a cacheable key and status list: check the signature, check the validity period, check revocation, and confirm the asset digest matches the content in question. The draft adds a rule for auditing the past that is thoughtful and slightly unusual, saying that a verifier evaluating a past use checks revocation status as of the time of use rather than as of now, since a grant may have been validly relied upon and revoked afterwards.

Its own security section then removes the reading most people would take from the word verifiable. A grant proves who issued it, when, and what was declared. It does not attest that the grantor holds the rights it purports to grant, and the draft spells out why in the sharpest sentence in the document: anyone in possession of an asset can compute its digest and issue a self-asserted grant over it. Establishing that an issuer is authoritative for an asset is explicitly out of scope. That is the same honest boundary an IETF compliance record draft drew two weeks ago, and the same one Web Bot Auth draws around what a signature actually proves: a signature establishes who is speaking, never whether they are telling the truth.

The draft is also unsure it belongs where it was filed. It notes that the AI Preferences working group has to date scoped mechanisms above the vocabulary, naming attachment, transmission, grants and identity, out of its charter, and suggests the DISPATCH area as a home instead while asking for confirmation. A mechanism nobody has agreed to host is a long way from anything a publisher can use, which is the same position pay per crawl occupies while it prices access with an undefined status code.

The verification procedure specified in section 5 of draft-wallace-aipref-grant-binding-01, with the as-of-use branch from section 6.1. Read from the draft on 15 August 2026; no implementation of this exists to measure.

The four page draft filed on the same day

The second submission of 6 August 2026 is draft-zehta-aipref-parameters-00, written by Timid Robot Zehta at Creative Commons, intended status informational, and four pages long including the boilerplate. Its abstract is one sentence: this document defines how parameters can be added to AI Preferences.

The move is small and costs nothing, which is the appeal. The vocabulary's serialization already relies on the Dictionary type from RFC 9651, Structured Field Values for HTTP, published in September 2024, and structured field items can already carry parameters, described there as an ordered map of key-value pairs associated with an item, with unique keys and values that are bare items and cannot themselves be parameterized. So the draft does not invent a syntax. It points at one that the chosen serialization already inherits and writes down the handling rules, of which the substantive one concerns URI references: a parameter whose value should be a URI reference and is not a valid one must be ignored, and a relative reference must be resolved before use.

Its three examples say more about intent than the prose does. The generic form is ai-train=n with a parameter foo=bar. The second is the interesting one, described as a content holder wanting to highlight the presence of a tip jar, written as ai-train=y with a tipjar parameter pointing at a path. The third converts categories from the display-based preferences draft into parameters, giving ai-train=n and search=y carrying display-text and a max-text-length of 160, which links this proposal to draft-madhavan-aipref-displaybasedpref-02 of 25 March 2026 and its attempt to control display rather than collection.

Notice what the tip jar example is doing. It is a permission with a condition attached, and the condition is a hint about compensation. That is a party-agnostic gesture toward the same problem the grant draft addresses head on: y and n cannot express a deal. Where the grant draft answers by building a second artifact held off the site, this one answers by making the on-site string carry more. It is the same instinct behind Earmark, which signs a preference into the content itself so it survives copying, and the three proposals disagree less about what is missing than about where to keep it.

Example stringBase preferenceParameterWhat it adds
ai-train=n;foo=barTraining disallowedfoo=barGeneric illustration of the syntax
ai-train=y;tipjar=/tipjarTraining allowedtipjar=/tipjarA path where a user could pay
ai-train=n,search=y;display-text=y;max-text-length=160Training disallowed, search alloweddisplay-text, max-text-lengthDisplay controls converted from a separate draft
The three worked examples from section 3.2 of draft-zehta-aipref-parameters-00, read on 15 August 2026, with what each expresses. None of these strings is deployable: the attachment mechanism that would put them on a site is not published.

None of this puts anything new on your site to fetch

The reason to cover a permission layer on a site that measures readability is that these two drafts point in opposite directions about where the answer lives, and only one of those directions leaves anything an external tool can observe.

The parameters draft keeps everything on the site: a longer string in whatever file eventually carries preferences, fetchable by anyone, checkable by anyone. The grant draft deliberately does the reverse. Its section 2 says so directly, describing preferences as attached to content and expressed to the world as a broadcast, party-agnostic signal, and a grant as a signed artifact bound to a named grantee, held and presented by that grantee, checkable without being attached to the content or announced to the world. If grants become how AI permissions actually work, then the public reservation on a site stops describing what is happening to it. A scan would keep reading the reservation correctly and keep being incomplete, in the same way a scan cannot see a Search Console setting or a Cloudflare rule. Saying so is not an argument against grants. It is the reason to say what a measurement covers, which is the whole content of our methodology page.

There is a second problem in the grant design that is specific to websites rather than to files, and it follows from the requirement that a grant must identify its asset by content digest. A digest binds a grant to one exact byte sequence. That is a good fit for a photograph, a dataset or a PDF, and an awkward one for a page that is generated per request. HTML changes on every deploy, and often between two consecutive fetches, which is not a hypothesis: the reason Lantad fetches a page twice, once raw and once through a browser, and compares the prose in each is that those two responses routinely differ, and prose parity is the name for the size of that difference. Which bytes a publisher would hash to grant rights over an article is not specified in the draft, and the answer is not obvious. Lantad has measured nothing about digest stability and holds no figure for how often a page's bytes change, so this is a reading of the specification rather than a finding.

What follows for a site owner today is short, because the honest answer is that nothing here is actionable yet. There is no file to write, no header to set, and no crawler that would read either. What remains worth doing is the ordinary work that decides whether any of this will ever matter on your domain: confirm what your robots.txt says to each named crawler with the robots.txt tester, look at what a crawler actually receives rather than what a browser renders, and check that the tokens you have written rules for are the ones currently in use in the crawler reference. A permission you cannot express is a smaller problem than an access you did not intend, and the second one is measurable now.

Sample Illustrative, not a measurement of any real site.

  • robots.txt rules Visible A public file at a fixed path, readable by any client without permission.
  • Preference string with parameters Visible if attached On-site by design, but the draft defining where to attach a preference is not published.
  • Grant credential Not visible Held by the grantee and checkable without being attached to the content or announced.
  • Whether a grant was revoked Not visible A status list the issuer controls, consulted by verifiers rather than published on the site.
  • Which bytes a digest covers Unspecified The draft requires a content digest and does not say how to compute one for a generated page.
What an external fetch of a site could observe under each proposal, following the distribution model each document describes. Illustrative of the mechanisms as specified, not a measurement of any real site.

Written by

Lantad

Published .

The web has a rich vocabulary for refusal and almost none for permission. A site can write a Disallow line, publish a machine readable reservation, or set a meta tag, and every one of those expresses the same shape of thing: no, or no to this crawler. Saying yes turns out to be much harder, because the useful version of yes is narrow. A publisher who has signed an agreement with one AI company wants to permit that company, over that content, until that date, with the ability to withdraw it later. Nothing a site can currently write down expresses any of those four things.

Common questions

Can an AI usage preference allow one company and disallow another?

No. The vocabulary at draft-ietf-aipref-vocab-06, read on 15 August 2026, assigns a value of y or n to a usage category and has no field for a recipient, so a statement applies to every party that reads it. Where two statements disagree, section 5.1 resolves to the most restrictive one, which means an allowance can never override a reservation on the same asset.

Does a robots.txt Allow line grant a crawler permission?

Not in the sense the word usually carries. RFC 9309 states that its rules are not a form of access authorization and that the protocol is not a substitute for valid content security measures. The user-agent line also addresses a product token, which section 2.2.1 describes as a name crawlers set for themselves, rather than a company.

Is the grant credential something a publisher can use now?

No. draft-wallace-aipref-grant-binding-01 is an individual submission dated 6 August 2026 that describes itself as a discussion starter and one candidate rather than a final design, and it is unsure of its own venue, noting that grants sit outside the AI Preferences working group charter and suggesting the DISPATCH area instead. No crawler vendor has said it would check one.

What can Lantad tell me about AI permissions on my site?

Only the part that is observable from outside: what your robots.txt says to each named crawler token, what status code and body a non-browser client receives, and how much of your prose is present before JavaScript runs. Lantad cannot see a preference held in a credential, a Search Console setting or a CDN rule, and reports what it did not measure rather than inferring it.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.