BlogFindings

Mobile first indexing and AI crawlers: 29 of 1,087 home pages answered a phone and a desktop differently

Google indexes the mobile version of a page and says so. The AI crawlers that decide whether ChatGPT and Perplexity can read a site publish user agent strings that name a desktop or no device at all. Lantad asked all 1,419 hostnames in this repository's committed corpus for their home page twice on 1 October 2026, once with a desktop platform token and once with a phone one, and 1,087 answered both. 989 of those returned an identical word count. 29 differed by a tenth or more, confirmed over two further rounds of fetching, and 26 of the 29 gave the phone less.

16 min read Lantad

So this run asked one narrow question of the whole corpus. Does the same URL return the same words to a request that says it comes from a desktop and a request that says it comes from a phone? The answer on 1 October 2026 is reassuring for most of the corpus and specific about where it is not. 989 of the 1,087 home pages that answered both requests returned an identical word count. 29 did not, and the work that matters in a measurement like this is not finding those 29 but proving they are real, because the first thing a two-fetch comparison picks up is pages that simply change between any two fetches for reasons that have nothing to do with who asked.

In short

  • Mobile first indexing is Google's rule and nobody else has adopted it in writing: Google's own documentation, Last updated 2025-12-10 UTC, says it uses the mobile version of a site's content crawled with the smartphone agent for indexing and ranking, while the six crawler user agent strings OpenAI and Perplexity published on 1 October 2026 name a Macintosh or no platform at all and never a phone.
  • Of the 1,087 corpus home pages that answered both a desktop-shaped and a phone-shaped request on 1 October 2026, 989 returned an identical word count and 854 were identical in byte length, so the ordinary case in this corpus is one document served to every device.
  • 29 home pages differed by a tenth or more in word count with the difference holding across three rounds of fetching on the same day, 26 of them giving the phone fewer words, and the gap reached 98 percent on alibaba.com, which answered the desktop request with 18,425 words and the phone with 362.
  • A word count gap is not the same as lost prose: of the 28 pages re-fetched for a vocabulary comparison, 16 lost a quarter or more of the distinct words the desktop page carried, while pophamdesign.com lost 6 of its 55 distinct words and all six were interface wording such as menu, close and navigate.
  • rbi.org.in sent the phone to m.rbi.org.in, whose home page carries a robots meta element reading noindex that the desktop page does not, which is the exact configuration Google's mobile first indexing guidance warns about when it says to use the same robots meta tags on the mobile and desktop site.
StageCountWhat it means
Hostnames asked1,419The committed corpus, an editorial frame rather than a random draw
Requests sent2,838Two per hostname, desktop platform token and phone platform token
Returned no HTTP status17886 connection failures, 62 read timeouts, 30 TLS failures
Answered HTTP 403402Refused the request outright, either agent
Answered 200 with HTML to both1,087The denominator for every comparison below
Answered 200 with HTML to the phone only18Before verification. Three still differed two rounds later
Answered 200 with HTML to the desktop only2Neither held up on re-fetch
Two GET requests for https://<host>/ per hostname, sent within a second of each other from one network location, redirects followed, 20 second timeout, no JavaScript executed. Both requests identified LantadBot and differed only in the platform token: one read Macintosh; Intel Mac OS X 10_15_7 and ended Safari/537.36, the other read Linux; Android 6.0.1; Nexus 5X and ended Mobile Safari/537.36. Measured by Lantad on 1 October 2026 across the 1,419 hostnames in this repository's two committed corpus seed files.

What mobile first indexing means, and which agent each engine sends

The rule is not folklore and it is not inferred. Google's mobile site and mobile first indexing best practices, carrying Last updated 2025-12-10 UTC when it was opened on 1 October 2026, opens by stating that Google uses the mobile version of a site's content, crawled with the smartphone agent, for indexing and ranking, and that this is called mobile first indexing. The same page names the three configurations a site can use. Responsive design serves the same HTML on the same URL to every device. Dynamic serving uses the same URL and, in the page's own words, relies on user agent sniffing and the Vary: user-agent HTTP response header to serve a different version of the HTML to different devices. Separate URLs serve different HTML on different addresses. The page then says its guidance only applies to the last two, because under responsive design the content and the metadata are the same either way.

Which agent does the asking is published too. Google's crawler reference, Last updated 2026-07-14 UTC, gives Googlebot Smartphone a user agent that opens Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) and carries the token Mobile Safari/537.36 before naming Googlebot. It lists a Googlebot Desktop string as well, so Google crawls with both and has said which one decides. That is the whole of the arrangement a site owner has been asked to plan for since mobile first indexing was announced, and it is a Google arrangement. No other engine has published anything equivalent, which is the gap this post is about rather than a complaint about Google. The vocabulary underneath all of it sits under AI crawler, and the tokens each vendor documents are listed in the crawler directory.

CrawlerVendorPlatform the string declares
Googlebot SmartphoneGoogleLinux; Android 6.0.1; Nexus 5X, ending Mobile Safari
Googlebot DesktopGoogleNo platform token, ending Safari/537.36
OAI-SearchBot/1.4OpenAIMacintosh; Intel Mac OS X 10_15_7
GPTBot/1.4OpenAINone at all
ChatGPT-User/1.0OpenAINone at all
OAI-AdsBot/1.0OpenAINone at all
PerplexityBot/1.0PerplexityNone at all
Perplexity-User/1.0PerplexityNone at all
Platform token in each crawler's own published user agent string. Google read at developers.google.com on 1 October 2026, page dated Last updated 2026-07-14 UTC. OpenAI read at developers.openai.com and Perplexity at docs.perplexity.ai on 1 October 2026; neither page carries a published revision date. A reading of vendor documentation, not a measurement of traffic.

Does mobile first indexing apply to AI crawlers?

Nothing read at source on 1 October 2026 says that it does, and nothing says that it does not. OpenAI's crawler overview describes four tokens and what each is for, publishes a user agent string for each, and says nothing about devices, viewports or which version of a page it expects. Of its four strings, one names a Macintosh and three name no platform at all. Perplexity's crawler page does the same for two tokens, both of which open Mozilla/5.0 AppleWebKit/537.36 with no platform in the parentheses. That is six published strings between two vendors and not one phone among them.

The registry this scanner ships is a useful cross-check because it is a transcription rather than a judgement. It holds 15 crawler tokens, stores the vendor's published user agent string for 11 of them, and exactly one of those 11 names a mobile platform: Bytespider, whose stored string opens Mozilla/5.0 (Linux; Android 5.0) and carries Mobile Safari/537.36. So the honest statement is narrow and it is enough to make the measurement worth taking. Most documented AI crawlers announce themselves as something other than a phone. A site that serves a different document to a phone is therefore serving its AI crawlers the version Google has said it does not index, and serving Google the version the AI crawlers do not see. Which way round that cuts depends entirely on which version is better, and that is a question about the site rather than about the crawler. This blog has measured the neighbouring question of which tokens are documented well enough to act on, and found that only six of fifteen crawler tokens publish a user agent to match, and that 1,004 robots.txt files name 2,209 distinct tokens between them.

  • Google States the rule Uses the mobile version crawled with the smartphone agent for indexing and ranking. Page dated Last updated 2025-12-10 UTC.
  • OpenAI Silent on device Four tokens documented with four user agent strings. One names a Macintosh, three name no platform. Devices are not mentioned.
  • Perplexity Silent on device Two tokens documented, both with no platform token in the user agent string. Devices are not mentioned.
  • What follows for a responsive site Nothing One document on one URL reaches every agent the same way, which is what 989 of 1,087 corpus home pages did.
  • What follows for a sniffing site Two audiences The version Google indexes and the version an AI crawler reads are different documents, and the site owner chose neither deliberately.
What each vendor's own documentation says about which version of a page it crawls, read at source on 1 October 2026. A reading of published pages, not a measurement.

989 of 1,087 home pages served a phone and a desktop the same words

This is the finding, and it is the opposite of alarming. Of the 1,087 hostnames that answered both requests with HTTP 200 and an HTML content type, 989 returned a word count that matched exactly, and 854 returned responses of identical byte length. 1,060 of the 1,087 carried a viewport meta element, which is the marker of a page built to lay itself out for whatever screen arrives rather than to be chosen between. Responsive design has won in this corpus, and for a site built that way the question this post asks has no answer worth having, because there is only one document.

Only 34 of the 1,087 differed in word count by a tenth or more before verification, 24 by a quarter or more and 15 by half or more. The header that is supposed to announce the difference is largely absent. The Vary response header exists to tell caches which request headers changed the response, and Google's own description of dynamic serving names Vary: user-agent as part of the configuration. 35 of the 1,087 desktop responses sent a Vary header naming User-Agent. Of the 29 differences that survived verification, 8 came with that header and 21 did not, so most of the sites varying their output by user agent are not declaring it. Nothing in this run measured what a cache or an intermediary did with that, and the count is simply what the responses carried. Word counts here mean text extracted from the raw bytes with script, style, template, noscript and SVG elements removed and no browser involved, which is the same reading used when this blog measured that JavaScript supplied 7.6 percent of the prose across 271 pages and the same notion of prose parity the scanner grades against.

  • Identical word count 989 sites 854 of these were identical in byte length too
  • Differed, under 10 percent 64 sites
  • Differed 10 to 24 percent 10 sites
  • Differed 25 to 49 percent 9 sites
  • Differed 50 percent or more 15 sites Before verification. Thirteen of the fifteen held
Agreement between the desktop-shaped and phone-shaped fetch, by band, across the 1,087 corpus home pages that answered HTTP 200 with HTML to both. Measured by Lantad on 1 October 2026.

Ruling out the page simply moving between two fetches

A difference between two requests is only a fact about user agents if the page is otherwise stable, and the cheapest way for this measurement to be wrong is for it to catch pages that change every time anybody loads them. So the 34 flagged hostnames, every one-sided reachability case, and a random control sample of 150 hostnames that showed no difference at all went through a second round: four fetches each, desktop, desktop, phone, desktop, with the phone request deliberately placed before a desktop one so that request order could not explain a gap.

The control answers the question the first pass cannot. Of the 229 hostnames in that round where both desktop fetches returned HTTP 200 and HTML, 223 returned an identical word count to two consecutive desktop requests, and only three moved by a tenth or more on their own: bestbuy.com, klarna.com and westjet.com. A page in this corpus is therefore stable on repeat, which is what makes the comparison readable. Applying that as a filter, 29 of the 34 flagged gaps held, meaning the three desktop fetches agreed within 5 percent of each other while the phone fetch sat a tenth or more away from every one of them. Five were rejected as the page moving rather than the agent mattering: admin.ch, bestbuy.com, jalan.net, klarna.com and westjet.com. A third round an hour later fetched the 29 twice with each agent again, and 28 reproduced the gap in the same direction. The exception is aliexpress.com, which on that round returned no extractable text to either agent, so it is counted in the 29 on the strength of two rounds and named here as the weakest of them. The method, the conduct of the crawler and what it declares are set out at methodology and the bot policy, and the 403 rate above is the same wall this blog hit when headless Chromium was blocked on 15 percent of sites.

The three rounds of fetching behind every confirmed figure in this post, all on 1 October 2026. D is a request carrying the desktop platform token, M a request carrying the phone platform token.

What the phone actually lost, and what was only interface wording

A word count is a blunt instrument and the gap it reports is not automatically a loss of prose. A page that repeats a navigation label twenty times on one layout and twice on another will show a large gap while saying the same things. So the 29 confirmed pages were fetched once more and compared by vocabulary rather than by volume: the set of distinct words of four letters or more on each version, and how much of the desktop page's vocabulary was absent from the phone's. 28 returned text to both agents on that attempt. 16 of the 28 lost a quarter or more of their distinct words and 10 lost half or more.

The two ends of that range are worth naming because they are different problems. At one end, pophamdesign.com showed a 47 percent word count gap and lost exactly 6 of its 55 distinct words, every one of them interface wording: back, close, items, menu, navigate and through. Nothing a reader or an answer engine wants was missing. At the other end, studio-chocolate.co.uk lost 113 of its 302 distinct words, among them bake, birthday, cake and corporate, which are the words a page about a chocolate studio exists to carry. spottedinprod.com lost 156 of 180, and the phone request also landed on a different path. dailymail.co.uk lost 1,986 distinct words of 3,011, cnn.com lost 671 of 1,083, and alibaba.com lost 988 of 1,083 while its word count fell from 18,425 to 362. These are pages where the version an undeclared crawler reads is a different article from the version a phone-shaped crawler reads, and the difference is measured in paragraphs rather than in furniture. The reading matters for the same reason 121,312 of 1,269,054 words sitting behind a marker mattered, and for the same reason 80,761 of 104,474 list items turned out to be navigation: what is in the bytes is not the same question as what is on the screen. Where a page offers a crawler nothing to follow at all, 84 of 1,091 home pages carried no internal link.

HostDesktop wordsPhone wordsDistinct words lost
alibaba.com18,425362988 of 1,083, or 91 percent
spottedinprod.com22681156 of 180, or 87 percent
carnival.com2193051 of 63, or 81 percent
practo.com579149153 of 217, or 71 percent
dailymail.co.uk11,6382,9091,986 of 3,011, or 66 percent
newegg.com2,069761283 of 434, or 65 percent
cnn.com3,1441,368671 of 1,083, or 62 percent
studio-chocolate.co.uk734420113 of 302, or 37 percent
pophamdesign.com1941026 of 55, all interface wording
The confirmed differences with the largest vocabulary loss. Word counts are from the first pass of 1 October 2026. Vocabulary loss is the share of the desktop page's distinct words of four letters or more absent from the phone version, measured on a separate fetch of the same 29 pages later the same day. No JavaScript was executed in either.

Three sites sent the phone somewhere else, and one put a noindex there

Three of the 1,087 used the separate URLs configuration Google's guidance describes, sending the phone request to a different hostname and keeping the desktop request where it was: rbi.org.in sent it to m.rbi.org.in, shein.com sent it from us.shein.com to m.shein.com, and taobao.com sent it from www.taobao.com to main.m.taobao.com. All three held across every round. That configuration is documented, it is allowed, and it is also the one where the two versions can drift apart without anybody noticing, because they are two documents maintained in two places.

One of the three shows exactly that drift. The home page at rbi.org.in carried no robots meta element on any desktop fetch. The page at m.rbi.org.in carried one reading noindex, on every attempt, which asks a crawler that honours it not to index the Reserve Bank of India's mobile home page. Google's mobile first indexing guidance addresses this case directly: it says to use the same robots meta tags on the mobile and desktop site, and that where a different tag is used on mobile, especially noindex or nofollow, Google may fail to crawl and index the page. Four other sites differed in that element: newegg.com sent index,follow,max-image-preview:large,max-snippet:-1 to the desktop and only index,follow to the phone, carnival.com sent follow to the desktop and nothing to the phone, aliexpress.com sent nothing to the desktop and follow,index to the phone, and rakuten.co.jp sent NOYDIR to the desktop and nothing to the phone. Markup moved too. newegg.com served two JSON-LD blocks to the desktop and none to the phone, while carnival.com served none to the desktop and one to the phone, both holding across rounds, and dailymail.co.uk went from five blocks to one. That is the same class of inconsistency as 163 of 1,069 home pages declaring no canonical tag, and it has the practical consequence described under structured data: an engine reading one version sees an entity the other version does not declare. Three of the 18 hostnames that answered only the phone in the first pass were still one-sided two rounds later. douglas.de and thl.fi answered HTTP 403 to every desktop fetch and HTTP 200 to the phone, and nykaa.com never completed a connection for a desktop fetch while answering the phone with HTTP 200 on three of four attempts. That is the inverse of the pattern this blog found when 103 of 1,089 sites served an unknown bot and refused GPTBot and when 79 of 115 refused a crawler their robots.txt allows. hilton.com looked like a third until the final round, when it refused both. What any given site should do about this is a question about its own two documents, and the fastest way to see which one a crawler gets is what GPTBot sees.

Desktop platform token

  • final URL: https://rbi.org.in/
  • title: Home | Official website of Reserve Bank of India
  • robots meta element: none present
  • words extracted: 1,492

Phone platform token

  • final URL: https://m.rbi.org.in//home.aspx
  • title: Reserve Bank of India
  • robots meta element: content reads noindex
  • words extracted: 2,247
rbi.org.in, the same request path taken by two agents, read at source on 1 October 2026 and repeated on every round. Strings as served, HTML entities decoded.

Written by

Lantad

Published .

Two different machines decide whether a page is readable, and they do not agree about what a page is. Google states plainly that it indexes the mobile version of a site. The crawlers that decide whether an answer engine can read the same site publish user agent strings that describe a desktop computer, or that describe no device at all. If one URL hands those two shapes of request two different documents, then search visibility and AI visibility are being decided from different text, and nothing on the page tells anybody that is happening.

Common questions

Does mobile first indexing apply to AI crawlers?

No published documentation read on 1 October 2026 says that it does. Mobile first indexing is Google's stated policy for Google Search: its guidance says Google uses the mobile version of a site's content, crawled with the smartphone agent, for indexing and ranking. OpenAI's and Perplexity's crawler pages document their tokens and user agent strings and say nothing about devices at all, and five of the six strings they publish name no platform or a Macintosh. So a site serving different content by device cannot assume either rule applies to the other engine.

Which version of my site does ChatGPT read?

Whichever version the site returns to a request carrying the user agent string OpenAI publishes, and those strings do not declare a phone. OAI-SearchBot's published string names Macintosh; Intel Mac OS X 10_15_7 and GPTBot's names no platform at all. On a responsive site the question does not arise, because there is one document. On a site that varies its HTML by user agent, the version returned to those strings is the version available to be cited, and this run found 29 corpus home pages where that version differs materially from the one a phone receives.

Is a word count gap between mobile and desktop a problem?

Not on its own, and this post separates the two cases rather than assuming. Of the 28 confirmed pages re-fetched for a vocabulary comparison on 1 October 2026, 16 lost a quarter or more of the distinct words the desktop version carried, which is prose going missing. pophamdesign.com showed a 47 percent gap in word count and lost only 6 distinct words, all of them interface wording such as menu and close, which is a layout difference and nothing more. The figure worth checking on a site is the vocabulary, not the volume.

What did this measurement not test?

It did not execute JavaScript, so content a site renders in the browser on one layout and not the other is invisible to it. It took one page per hostname, always the home page, from one network location, so nothing here describes interior pages or what a site returns elsewhere in the world. It sent a user agent naming LantadBot with a borrowed platform token rather than any real crawler's string, so it measures how a site responds to the platform token and not how it responds to GPTBot or Googlebot specifically. It also observed no crawler reading anything: no server logs were held, and no figure here says what any answer engine did with either version.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.