Blog / Time is selling sponsored blocks inside the markdown AI crawlers read

Time is selling sponsored blocks inside the markdown AI crawlers read

Digiday reported on 30 July 2026 that Time has converted its pages to markdown for AI systems and is selling FAQ shaped sponsored content inside those files, which turns the machine copy of a page into something that can differ from the human one by design.

In short

  • Digiday reported on 30 July 2026, in a piece by Sara Guaglione, that Time has begun serving sponsored FAQ blocks inside markdown versions of its articles, produced with the advertising technology company Mobian, and named Ally Bank and the Project Management Institute as the first two sponsors.
  • A markdown twin is a second representation of the same article, so any site serving one now holds two copies of its content that are free to disagree, and only one of the two is the copy a human reader ever opens.
  • The llms.txt proposal published by Jeremy Howard on 3 September 2024 asks sites to serve a clean markdown version of a page at the same URL with .md appended, which is the convention lantad.co follows on every public page.
  • Google's search spam policies, carrying Last updated 2026-05-15 UTC, define cloaking as presenting different content to users and search engines with intent to manipulate rankings, and give inserting text only when the requesting user agent is a search engine as an example.
  • Lantad did not fetch Time's markdown files, did not establish what URLs they are served at, and scores no signal for sponsored disclosure, so nothing here is a measurement of Time's site.

For two years the argument about what an AI crawler should be allowed to take has been an argument about access: which tokens to disallow, which files to write, which requests to let through. Markdown for AI crawlers changes the question, because it is not about whether the crawler gets the page. It is about which page it gets. A site that publishes a markdown twin has decided to hand machines a different artefact from the one it hands people, and once that artefact exists somebody will eventually put something in it that is not in the other one.

Somebody now has. Lantad measured none of what follows and did not fetch a single file from the publisher in question. This is a reading of trade reporting and of two published specifications, and the part this site can speak to arrives after them: a second representation is exactly the condition prose parity was built to detect, and the scanner that detects it has no way of knowing whether a difference was an accident or a sale.

What the reporting establishes

  • Time serves markdown versions of articles
  • Text only, without design or images
  • Sponsored blocks shaped as FAQs sit inside them
  • Labelled as sponsored content
  • Produced with the ad technology firm Mobian
  • Ally Bank and the Project Management Institute are the first sponsors
  • Conversion to markdown began the previous month

What it does not establish

  • The URLs the markdown files are served at
  • Whether the HTML article carries the same block
  • Whether any AI system weights the block at all
  • How many articles have been converted
  • Whether any engine discounts sponsored text
  • Anything about publishers other than the one named
Reported by Digiday on 30 July 2026 in a piece by Sara Guaglione. A restatement of published reporting, not a measurement of any site.

What was reported, and what the reporting actually says

The source is Digiday, which published a piece by Sara Guaglione on 30 July 2026 headed Time has started serving ads to AI agents, at digiday.com/media/time-has-started-serving-ads-to-ai-agents/. That URL is written here as plain text rather than as a link because this site keeps a registered list of the external hosts it will link to, and a trade publication is something to cite by name rather than to hand link authority to. The reporting is the primary account and everything in this section comes from it.

What it describes is two separate decisions that arrived together, and they are worth pulling apart because only the second one is new. The first is that Time converted its webpages into markdown versions, described as stripped down text only copies without design or images, and began doing so in the month before the report. That decision on its own is unremarkable and increasingly common. The second is that Time has begun placing sponsored content inside those markdown files, formatted as FAQ entries carrying brand information and labelled as sponsored, produced in partnership with an advertising technology company called Mobian, with Ally Bank and the Project Management Institute named as the first two sponsors.

The rationale given is traffic. Time's chief operating officer Mark Howard is quoted describing the agent audience as a growing traffic source and therefore a growing source of inventory, and the piece reports that Time sees more bot traffic than human traffic on most days. Whether that inventory has any value depends entirely on a question nobody in the piece claims to have answered, which is whether the systems doing the fetching treat a labelled sponsored FAQ as content worth repeating or as promotional material worth discounting.

That uncertainty is the honest headline, and it is worth holding onto while the rest of the industry writes this up as a new advertising channel. Nothing in the reporting establishes an effect. It establishes that a large publisher has built the inventory and found two buyers, which is a fact about a business decision rather than a fact about how models behave. It also lands awkwardly beside Google's own position: the same week, this blog read Google's guide to generative AI features, whose Mythbusting section states you do not need to create Markdown to appear in Google Search because Google Search does not use it. Google is not the buyer here. The audience for a markdown twin is the set of systems that do fetch it, and ChatGPT and its peers publish nothing comparable about how they weight one. That gap between a documented position and an undocumented one is most of what makes generative engine optimization hard to do honestly.

ReportedAttributionWhat it does not show
Pages converted to markdown, text only, no design or imagesDigiday, 30 July 2026The URL scheme, or how many articles
Sponsored blocks formatted as FAQs, labelled as sponsoredDigiday, 30 July 2026Whether the same block appears in the HTML
Produced with the ad technology company MobianDigiday, 30 July 2026How the blocks are selected per article
Ally Bank and the Project Management Institute as first sponsorsDigiday, 30 July 2026Any spend, term or performance figure
Time sees more bot traffic than human traffic most daysTime's COO, quoted by DigidayA measured share, which the piece does not give
Conversion began the month before publicationDigiday, 30 July 2026An exact start date
Claims in the Digiday report of 30 July 2026, and the scope each one carries. Every row restates published reporting rather than an observation of any file.

A markdown twin is a second copy, and two copies can disagree

The convention Time appears to be following is not new and did not come from a publisher. It comes from the llms.txt proposal, published by Jeremy Howard on 3 September 2024, which does two things at once. The better known half asks for an index file at the site root. The half that matters here is the second: the llms.txt proposal states that pages carrying information useful for language models should provide a clean markdown version of those pages at the same URL as the original with .md appended, with index.html.md for URLs that have no file name.

This site does that. Every public page on lantad.co is served as markdown at its own URL plus .md, indexed at the markdown index, and the build refuses to pass if a public page exists without a markdown twin. That is not an aside offered for balance. It means the thing described in this post is a thing this site does, minus the advertising, and any argument made below about the risks of a second copy is an argument that applies here first.

It also means the failure mode is one worth naming precisely. The twin is generated from the same typed data the page renders from, so the two cannot drift apart without a code change. That is a design choice with a cost and a benefit: nothing can be written into the markdown that is not on the page, which also means nothing can be sold there. A publisher generating markdown through a separate pipeline, from a separate vendor, has no such constraint, and there is no reason it would want one if the pipeline exists in order to insert something.

None of which says the file is worth serving. This site has published, at length, that the evidence for llms.txt is poor: what the evidence says about llms.txt reported third party measurement showing the overwhelming majority of valid files are never fetched at all, and it was published while this site was shipping an llms.txt generator and maintaining a glossary entry for the file. The position has not changed. Serving a machine readable copy is cheap and defensible on grounds of structure rather than ranking, and anyone selling it as a visibility lever is selling something no published measurement supports. The new development is not that the file works. It is that a second copy is now somewhere revenue can be placed, which gives the copy a reason to exist independent of whether any model reads it.

Sample Illustrative, not a measurement of any real site.

One article, two representations

  • GET /news/example-article 200 text/html, full page with navigation, markup and images
  • GET /news/example-article.md 200 text/markdown, prose only, structure explicit
  • Compare body text of the two responses The only step that establishes whether they agree
  • If the markdown carries a block the HTML does not A difference exists; the reason for it is not in the response
  • If both carry the same block One document served in two formats, which is the intended case
The two requests a crawler can make for one article under the .md convention described in the llms.txt proposal. Illustrative of the mechanism, not a capture of any real site.

Prose parity is the thing a second representation puts at risk

Lantad's score is four components and the largest of them is parity, weighted at 50 points of 100 in core/src/config.ts against 25 for access, 15 for structure and 10 for schema. That figure is a decision rather than a discovery, and the distinction matters enough to repeat: nobody measured that parity is worth half of anything. Somebody chose the weight, it is recorded in a config file, and the methodology page is where the choice is written down and defended rather than hidden inside a number.

What the component measures is narrow and it is worth being exact about it, because the word parity gets used loosely. It compares the prose a crawler receives against the prose a browser receives for the same URL, which is a test built for the ordinary modern failure: a page whose text arrives only after JavaScript executes, so a client that does not execute JavaScript receives a shell. That failure is accidental. Nobody ships a single page application in order to hide their article, and the six captures behind what a crawler meets on a real storefront are a record of that accident happening at different severities on stores running the same platform.

A markdown twin is a different shape of problem and the comparison does not straightforwardly cover it. The twin lives at its own URL, so a scan of the HTML page and a scan of the .md file are two scans of two documents, each internally consistent. Whatever difference exists between them sits in the gap between two URLs rather than between two clients of one URL, and a checker built for the second gap will report both documents as fine. What GPTBot sees fetches a page as a crawler would and shows the text that survives; it does not go looking for a parallel document to diff against, and it would be dishonest to imply otherwise.

Whether the distinction matters to a model depends on which document gets fetched, and that is not uniform. Common Crawl's July archive was collected by a crawler executing no JavaScript, so for that consumer the rendered page was never the artefact anyway. Retrieval crawlers that run at answer time behave differently again. The result is that a publisher with two documents has two audiences, split by which client asked, and no way to reconcile what each was told other than keeping the documents the same.

  • Prose parity 50 pts Crawler text against browser text for one URL. Built for the JavaScript shell case, not for a parallel document at another URL.
  • Access 25 pts Read from robots.txt and response headers, so it survives a failed body fetch when the other three cannot.
  • Structure 15 pts Headings, semantics and machine readable extras. The llms.txt presence check sits here at half the weight of a full check.
  • Schema 10 pts Structured data on the page. The smallest of the four, and unrelated to which representation was served.
The four score components and their weights, read from SCORE_WEIGHTS in core/src/config.ts. These are configured settings, not measured findings about the web.

Where the cloaking line sits, and why this is not obviously over it

The word hanging over all of this is cloaking, and it should be used carefully or not at all. Google's search spam policies, carrying Last updated 2026-05-15 UTC when read on 1 August 2026, define it as the practice of presenting different content to users and search engines with the intent to manipulate search rankings and mislead users. Two of the elements in that sentence do work that summaries usually drop. It requires a difference between what users and search engines receive, and it requires intent to manipulate rankings and mislead.

The examples the policy gives are more specific still. One is showing a page about travel destinations to search engines while showing a page about discount drugs to users. The other, closer to the subject here, is inserting text or keywords into a page only when the user agent requesting the page is a search engine rather than a human visitor. That second example is about conditional serving: the same URL returning different bodies depending on who asked.

A markdown twin at a separate URL is not that, at least not on its face. Anyone can fetch the .md file, no user agent test is involved, and a labelled sponsored block is disclosed rather than concealed. It is also worth saying plainly that Google's policy governs Google Search, and Google Search is not the consumer a markdown twin is built for. Applying a Google spam definition to a file Google says it does not read is a category error, and this post is not making the accusation.

The uncomfortable part is what the two cases have in common rather than where they differ. Both produce a body of text that shapes what a machine says about a subject and that no human reader is routed to. The disclosure that makes the sponsored block honest is a label inside a document whose entire purpose is to be consumed by something that does not read labels the way a person does, and no vendor has published how, or whether, its retrieval treats a sponsorship marker. That is not an accusation either. It is an admission that the mechanism that makes disclosure work for humans has no documented equivalent here, and it is the reason this site refuses to convert an unmeasured condition into a grade, which is the argument in why we withhold a grade. Where a claim is genuinely testable, such as which crawlers a robots.txt admits before any of this becomes relevant, the answer is available; how to get cited by Claude starts there rather than with content strategy for the same reason.

Sample Illustrative, not a measurement of any real site.

  • Same text, two formats Intended The markdown twin is generated from the same source as the page, so neither can carry anything the other lacks.
  • Markdown carries an extra block A difference Two documents at two URLs, both fetchable by anyone. The difference is real and the reason for it is not in the response.
  • Text inserted only for a crawler user agent The policy example One URL returning different bodies by requester. This is the case Google's spam policies name directly.
  • Rendered page is empty without JavaScript The ordinary failure Nobody intended it, which is why parity is scored as a defect rather than as an accusation.
Four ways one article can differ between what a machine receives and what a person receives. Constructed to separate the cases, not observed on any site.

What to check on your own site if you serve markdown to crawlers

None of this requires a scanner, and the checks below are all things a site owner can run with a terminal in a few minutes. The first is simply whether a twin exists at all. Append .md to one of your article URLs and fetch it. A 404 means the convention is not deployed and nothing else in this post applies to you. A 200 means you have a second document, and the only question that follows is whether you know what is in it.

The second check is a diff, and it is the one worth doing properly. Fetch both representations of the same article and compare the body text, not the byte count. What you are looking for is any block present in one and absent from the other: a promotional insert, a boilerplate footer, a summary that was generated rather than written, or a paragraph that a conversion pipeline dropped. If a third party generates your markdown, this diff is the only thing standing between your editorial voice and whatever that pipeline decides belongs in it.

The third is provenance. Find out which system produces the markdown and whether it shares a source with the page. Generated from the same data means the two cannot silently diverge. Generated by a separate service means they can, and the question of who reviews the output before it ships is a real editorial question rather than a technical one.

The fourth is access, because none of the above matters if the files are not reachable. The AI crawlers tool lists the tokens worth thinking about and the robots.txt tester evaluates a file per crawler rather than in the aggregate, which is the distinction most robots checkers lose. Confirm your markdown paths are not caught by a Disallow rule written for something else, and remember that robots.txt is only the outer of the two gates: two layers decide whether AI can read your site covers the edge and network layer that sits above it and that your own file will never mention.

The last one is a decision rather than a check. If you are going to publish a document written for machines, decide now whether anything may appear in it that does not appear on the page, and write the answer down somewhere a future pipeline change has to pass. Time has answered that question one way and disclosed it. The failure worth avoiding is not answering it at all and discovering the answer later, in a file nobody on the editorial side has read.

The order to check a markdown twin in, from cheapest to most consequential. A procedure, not a measurement.

Related

Common questions

What did Time actually start doing with markdown?

According to Digiday's report of 30 July 2026 by Sara Guaglione, Time converted its webpages into markdown versions, described as text only copies without design or images, and began placing sponsored content inside them formatted as FAQ entries and labelled as sponsored. The work was done with an advertising technology company called Mobian, and Ally Bank and the Project Management Institute were named as the first sponsors. Lantad did not fetch any of those files.

Is serving a markdown version of a page cloaking?

Not on the definition Google publishes. Google's search spam policies, carrying Last updated 2026-05-15 UTC, define cloaking as presenting different content to users and search engines with the intent to manipulate search rankings and mislead users, and the example closest to this case is inserting text only when the requesting user agent is a search engine. A markdown file at its own URL is fetchable by anyone and involves no user agent test. The open question is not whether it meets that definition but whether the two documents say the same thing.

Does lantad.co serve markdown versions of its own pages?

Yes. Every public page is served as markdown at the same URL with .md appended, following the convention in the llms.txt proposal published on 3 September 2024, and the index is at /md. The files are generated from the same typed data the pages render from, so neither can carry content the other lacks, and the build fails if a public page exists without a markdown twin. No sponsored content appears in any of them.

Does Lantad's parity check catch a difference between a page and its markdown twin?

No, and it is not built to. Prose parity compares the text a crawler receives against the text a browser receives for one URL, which catches the common case of a page whose content arrives only after JavaScript runs. A markdown twin lives at a separate URL, so both documents can be internally consistent while differing from each other. Diffing the two is a manual check today.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.