BlogFindings
Mobile first indexing and AI crawlers: 29 of 1,087 home pages answered a phone and a desktop differently
Google indexes the mobile version of a page and says so. The AI crawlers that decide whether ChatGPT and Perplexity can read a site publish user agent strings that name a desktop or no device at all. Lantad asked all 1,419 hostnames in this repository's committed corpus for their home page twice on 1 October 2026, once with a desktop platform token and once with a phone one, and 1,087 answered both. 989 of those returned an identical word count. 29 differed by a tenth or more, confirmed over two further rounds of fetching, and 26 of the 29 gave the phone less.
So this run asked one narrow question of the whole corpus. Does the same URL return the same words to a request that says it comes from a desktop and a request that says it comes from a phone? The answer on 1 October 2026 is reassuring for most of the corpus and specific about where it is not. 989 of the 1,087 home pages that answered both requests returned an identical word count. 29 did not, and the work that matters in a measurement like this is not finding those 29 but proving they are real, because the first thing a two-fetch comparison picks up is pages that simply change between any two fetches for reasons that have nothing to do with who asked.
In short
- Mobile first indexing is Google's rule and nobody else has adopted it in writing: Google's own documentation, Last updated 2025-12-10 UTC, says it uses the mobile version of a site's content crawled with the smartphone agent for indexing and ranking, while the six crawler user agent strings OpenAI and Perplexity published on 1 October 2026 name a Macintosh or no platform at all and never a phone.
- Of the 1,087 corpus home pages that answered both a desktop-shaped and a phone-shaped request on 1 October 2026, 989 returned an identical word count and 854 were identical in byte length, so the ordinary case in this corpus is one document served to every device.
- 29 home pages differed by a tenth or more in word count with the difference holding across three rounds of fetching on the same day, 26 of them giving the phone fewer words, and the gap reached 98 percent on alibaba.com, which answered the desktop request with 18,425 words and the phone with 362.
- A word count gap is not the same as lost prose: of the 28 pages re-fetched for a vocabulary comparison, 16 lost a quarter or more of the distinct words the desktop page carried, while pophamdesign.com lost 6 of its 55 distinct words and all six were interface wording such as menu, close and navigate.
- rbi.org.in sent the phone to m.rbi.org.in, whose home page carries a robots meta element reading noindex that the desktop page does not, which is the exact configuration Google's mobile first indexing guidance warns about when it says to use the same robots meta tags on the mobile and desktop site.
| Stage | Count | What it means |
|---|---|---|
| Hostnames asked | 1,419 | The committed corpus, an editorial frame rather than a random draw |
| Requests sent | 2,838 | Two per hostname, desktop platform token and phone platform token |
| Returned no HTTP status | 178 | 86 connection failures, 62 read timeouts, 30 TLS failures |
| Answered HTTP 403 | 402 | Refused the request outright, either agent |
| Answered 200 with HTML to both | 1,087 | The denominator for every comparison below |
| Answered 200 with HTML to the phone only | 18 | Before verification. Three still differed two rounds later |
| Answered 200 with HTML to the desktop only | 2 | Neither held up on re-fetch |
What mobile first indexing means, and which agent each engine sends
The rule is not folklore and it is not inferred. Google's mobile site and mobile first indexing best practices, carrying Last updated 2025-12-10 UTC when it was opened on 1 October 2026, opens by stating that Google uses the mobile version of a site's content, crawled with the smartphone agent, for indexing and ranking, and that this is called mobile first indexing. The same page names the three configurations a site can use. Responsive design serves the same HTML on the same URL to every device. Dynamic serving uses the same URL and, in the page's own words, relies on user agent sniffing and the Vary: user-agent HTTP response header to serve a different version of the HTML to different devices. Separate URLs serve different HTML on different addresses. The page then says its guidance only applies to the last two, because under responsive design the content and the metadata are the same either way.
Which agent does the asking is published too. Google's crawler reference, Last updated 2026-07-14 UTC, gives Googlebot Smartphone a user agent that opens Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) and carries the token Mobile Safari/537.36 before naming Googlebot. It lists a Googlebot Desktop string as well, so Google crawls with both and has said which one decides. That is the whole of the arrangement a site owner has been asked to plan for since mobile first indexing was announced, and it is a Google arrangement. No other engine has published anything equivalent, which is the gap this post is about rather than a complaint about Google. The vocabulary underneath all of it sits under AI crawler, and the tokens each vendor documents are listed in the crawler directory.
| Crawler | Vendor | Platform the string declares |
|---|---|---|
| Googlebot Smartphone | Linux; Android 6.0.1; Nexus 5X, ending Mobile Safari | |
| Googlebot Desktop | No platform token, ending Safari/537.36 | |
| OAI-SearchBot/1.4 | OpenAI | Macintosh; Intel Mac OS X 10_15_7 |
| GPTBot/1.4 | OpenAI | None at all |
| ChatGPT-User/1.0 | OpenAI | None at all |
| OAI-AdsBot/1.0 | OpenAI | None at all |
| PerplexityBot/1.0 | Perplexity | None at all |
| Perplexity-User/1.0 | Perplexity | None at all |
Does mobile first indexing apply to AI crawlers?
Nothing read at source on 1 October 2026 says that it does, and nothing says that it does not. OpenAI's crawler overview describes four tokens and what each is for, publishes a user agent string for each, and says nothing about devices, viewports or which version of a page it expects. Of its four strings, one names a Macintosh and three name no platform at all. Perplexity's crawler page does the same for two tokens, both of which open Mozilla/5.0 AppleWebKit/537.36 with no platform in the parentheses. That is six published strings between two vendors and not one phone among them.
The registry this scanner ships is a useful cross-check because it is a transcription rather than a judgement. It holds 15 crawler tokens, stores the vendor's published user agent string for 11 of them, and exactly one of those 11 names a mobile platform: Bytespider, whose stored string opens Mozilla/5.0 (Linux; Android 5.0) and carries Mobile Safari/537.36. So the honest statement is narrow and it is enough to make the measurement worth taking. Most documented AI crawlers announce themselves as something other than a phone. A site that serves a different document to a phone is therefore serving its AI crawlers the version Google has said it does not index, and serving Google the version the AI crawlers do not see. Which way round that cuts depends entirely on which version is better, and that is a question about the site rather than about the crawler. This blog has measured the neighbouring question of which tokens are documented well enough to act on, and found that only six of fifteen crawler tokens publish a user agent to match, and that 1,004 robots.txt files name 2,209 distinct tokens between them.
-
GoogleStates the rule Uses the mobile version crawled with the smartphone agent for indexing and ranking. Page dated Last updated 2025-12-10 UTC. -
OpenAISilent on device Four tokens documented with four user agent strings. One names a Macintosh, three name no platform. Devices are not mentioned. -
PerplexitySilent on device Two tokens documented, both with no platform token in the user agent string. Devices are not mentioned. -
What follows for a responsive siteNothing One document on one URL reaches every agent the same way, which is what 989 of 1,087 corpus home pages did. -
What follows for a sniffing siteTwo audiences The version Google indexes and the version an AI crawler reads are different documents, and the site owner chose neither deliberately.
989 of 1,087 home pages served a phone and a desktop the same words
This is the finding, and it is the opposite of alarming. Of the 1,087 hostnames that answered both requests with HTTP 200 and an HTML content type, 989 returned a word count that matched exactly, and 854 returned responses of identical byte length. 1,060 of the 1,087 carried a viewport meta element, which is the marker of a page built to lay itself out for whatever screen arrives rather than to be chosen between. Responsive design has won in this corpus, and for a site built that way the question this post asks has no answer worth having, because there is only one document.
Only 34 of the 1,087 differed in word count by a tenth or more before verification, 24 by a quarter or more and 15 by half or more. The header that is supposed to announce the difference is largely absent. The Vary response header exists to tell caches which request headers changed the response, and Google's own description of dynamic serving names Vary: user-agent as part of the configuration. 35 of the 1,087 desktop responses sent a Vary header naming User-Agent. Of the 29 differences that survived verification, 8 came with that header and 21 did not, so most of the sites varying their output by user agent are not declaring it. Nothing in this run measured what a cache or an intermediary did with that, and the count is simply what the responses carried. Word counts here mean text extracted from the raw bytes with script, style, template, noscript and SVG elements removed and no browser involved, which is the same reading used when this blog measured that JavaScript supplied 7.6 percent of the prose across 271 pages and the same notion of prose parity the scanner grades against.
Ruling out the page simply moving between two fetches
A difference between two requests is only a fact about user agents if the page is otherwise stable, and the cheapest way for this measurement to be wrong is for it to catch pages that change every time anybody loads them. So the 34 flagged hostnames, every one-sided reachability case, and a random control sample of 150 hostnames that showed no difference at all went through a second round: four fetches each, desktop, desktop, phone, desktop, with the phone request deliberately placed before a desktop one so that request order could not explain a gap.
The control answers the question the first pass cannot. Of the 229 hostnames in that round where both desktop fetches returned HTTP 200 and HTML, 223 returned an identical word count to two consecutive desktop requests, and only three moved by a tenth or more on their own: bestbuy.com, klarna.com and westjet.com. A page in this corpus is therefore stable on repeat, which is what makes the comparison readable. Applying that as a filter, 29 of the 34 flagged gaps held, meaning the three desktop fetches agreed within 5 percent of each other while the phone fetch sat a tenth or more away from every one of them. Five were rejected as the page moving rather than the agent mattering: admin.ch, bestbuy.com, jalan.net, klarna.com and westjet.com. A third round an hour later fetched the 29 twice with each agent again, and 28 reproduced the gap in the same direction. The exception is aliexpress.com, which on that round returned no extractable text to either agent, so it is counted in the 29 on the strength of two rounds and named here as the weakest of them. The method, the conduct of the crawler and what it declares are set out at methodology and the bot policy, and the 403 rate above is the same wall this blog hit when headless Chromium was blocked on 15 percent of sites.
Flow: Round 1: D then M, 1,419 hostnames to 34 word gaps, 20 one-sided; 34 word gaps, 20 one-sided to Round 2: D, D, M, D on 246 hostnames; Round 2: D, D, M, D on 246 hostnames to Control: 223 of 229 stable on repeat; Control: 223 of 229 stable on repeat to 29 gaps held, 5 rejected; 29 gaps held, 5 rejected to Round 3: D, D, M, M an hour later; Round 3: D, D, M, M an hour later to 28 of 29 reproduced.
What the phone actually lost, and what was only interface wording
A word count is a blunt instrument and the gap it reports is not automatically a loss of prose. A page that repeats a navigation label twenty times on one layout and twice on another will show a large gap while saying the same things. So the 29 confirmed pages were fetched once more and compared by vocabulary rather than by volume: the set of distinct words of four letters or more on each version, and how much of the desktop page's vocabulary was absent from the phone's. 28 returned text to both agents on that attempt. 16 of the 28 lost a quarter or more of their distinct words and 10 lost half or more.
The two ends of that range are worth naming because they are different problems. At one end, pophamdesign.com showed a 47 percent word count gap and lost exactly 6 of its 55 distinct words, every one of them interface wording: back, close, items, menu, navigate and through. Nothing a reader or an answer engine wants was missing. At the other end, studio-chocolate.co.uk lost 113 of its 302 distinct words, among them bake, birthday, cake and corporate, which are the words a page about a chocolate studio exists to carry. spottedinprod.com lost 156 of 180, and the phone request also landed on a different path. dailymail.co.uk lost 1,986 distinct words of 3,011, cnn.com lost 671 of 1,083, and alibaba.com lost 988 of 1,083 while its word count fell from 18,425 to 362. These are pages where the version an undeclared crawler reads is a different article from the version a phone-shaped crawler reads, and the difference is measured in paragraphs rather than in furniture. The reading matters for the same reason 121,312 of 1,269,054 words sitting behind a marker mattered, and for the same reason 80,761 of 104,474 list items turned out to be navigation: what is in the bytes is not the same question as what is on the screen. Where a page offers a crawler nothing to follow at all, 84 of 1,091 home pages carried no internal link.
| Host | Desktop words | Phone words | Distinct words lost |
|---|---|---|---|
| alibaba.com | 18,425 | 362 | 988 of 1,083, or 91 percent |
| spottedinprod.com | 226 | 81 | 156 of 180, or 87 percent |
| carnival.com | 219 | 30 | 51 of 63, or 81 percent |
| practo.com | 579 | 149 | 153 of 217, or 71 percent |
| dailymail.co.uk | 11,638 | 2,909 | 1,986 of 3,011, or 66 percent |
| newegg.com | 2,069 | 761 | 283 of 434, or 65 percent |
| cnn.com | 3,144 | 1,368 | 671 of 1,083, or 62 percent |
| studio-chocolate.co.uk | 734 | 420 | 113 of 302, or 37 percent |
| pophamdesign.com | 194 | 102 | 6 of 55, all interface wording |
Three sites sent the phone somewhere else, and one put a noindex there
Three of the 1,087 used the separate URLs configuration Google's guidance describes, sending the phone request to a different hostname and keeping the desktop request where it was: rbi.org.in sent it to m.rbi.org.in, shein.com sent it from us.shein.com to m.shein.com, and taobao.com sent it from www.taobao.com to main.m.taobao.com. All three held across every round. That configuration is documented, it is allowed, and it is also the one where the two versions can drift apart without anybody noticing, because they are two documents maintained in two places.
One of the three shows exactly that drift. The home page at rbi.org.in carried no robots meta element on any desktop fetch. The page at m.rbi.org.in carried one reading noindex, on every attempt, which asks a crawler that honours it not to index the Reserve Bank of India's mobile home page. Google's mobile first indexing guidance addresses this case directly: it says to use the same robots meta tags on the mobile and desktop site, and that where a different tag is used on mobile, especially noindex or nofollow, Google may fail to crawl and index the page. Four other sites differed in that element: newegg.com sent index,follow,max-image-preview:large,max-snippet:-1 to the desktop and only index,follow to the phone, carnival.com sent follow to the desktop and nothing to the phone, aliexpress.com sent nothing to the desktop and follow,index to the phone, and rakuten.co.jp sent NOYDIR to the desktop and nothing to the phone. Markup moved too. newegg.com served two JSON-LD blocks to the desktop and none to the phone, while carnival.com served none to the desktop and one to the phone, both holding across rounds, and dailymail.co.uk went from five blocks to one. That is the same class of inconsistency as 163 of 1,069 home pages declaring no canonical tag, and it has the practical consequence described under structured data: an engine reading one version sees an entity the other version does not declare. Three of the 18 hostnames that answered only the phone in the first pass were still one-sided two rounds later. douglas.de and thl.fi answered HTTP 403 to every desktop fetch and HTTP 200 to the phone, and nykaa.com never completed a connection for a desktop fetch while answering the phone with HTTP 200 on three of four attempts. That is the inverse of the pattern this blog found when 103 of 1,089 sites served an unknown bot and refused GPTBot and when 79 of 115 refused a crawler their robots.txt allows. hilton.com looked like a third until the final round, when it refused both. What any given site should do about this is a question about its own two documents, and the fastest way to see which one a crawler gets is what GPTBot sees.
Desktop platform token
- final URL: https://rbi.org.in/
- title: Home | Official website of Reserve Bank of India
- robots meta element: none present
- words extracted: 1,492
Phone platform token
- final URL: https://m.rbi.org.in//home.aspx
- title: Reserve Bank of India
- robots meta element: content reads noindex
- words extracted: 2,247
Lantad
Published .
Two different machines decide whether a page is readable, and they do not agree about what a page is. Google states plainly that it indexes the mobile version of a site. The crawlers that decide whether an answer engine can read the same site publish user agent strings that describe a desktop computer, or that describe no device at all. If one URL hands those two shapes of request two different documents, then search visibility and AI visibility are being decided from different text, and nothing on the page tells anybody that is happening.
Common questions
Does mobile first indexing apply to AI crawlers?
No published documentation read on 1 October 2026 says that it does. Mobile first indexing is Google's stated policy for Google Search: its guidance says Google uses the mobile version of a site's content, crawled with the smartphone agent, for indexing and ranking. OpenAI's and Perplexity's crawler pages document their tokens and user agent strings and say nothing about devices at all, and five of the six strings they publish name no platform or a Macintosh. So a site serving different content by device cannot assume either rule applies to the other engine.
Which version of my site does ChatGPT read?
Whichever version the site returns to a request carrying the user agent string OpenAI publishes, and those strings do not declare a phone. OAI-SearchBot's published string names Macintosh; Intel Mac OS X 10_15_7 and GPTBot's names no platform at all. On a responsive site the question does not arise, because there is one document. On a site that varies its HTML by user agent, the version returned to those strings is the version available to be cited, and this run found 29 corpus home pages where that version differs materially from the one a phone receives.
Is a word count gap between mobile and desktop a problem?
Not on its own, and this post separates the two cases rather than assuming. Of the 28 confirmed pages re-fetched for a vocabulary comparison on 1 October 2026, 16 lost a quarter or more of the distinct words the desktop version carried, which is prose going missing. pophamdesign.com showed a 47 percent gap in word count and lost only 6 distinct words, all of them interface wording such as menu and close, which is a layout difference and nothing more. The figure worth checking on a site is the vocabulary, not the volume.
What did this measurement not test?
It did not execute JavaScript, so content a site renders in the browser on one layout and not the other is invisible to it. It took one page per hostname, always the home page, from one network location, so nothing here describes interior pages or what a site returns elsewhere in the world. It sent a user agent naming LantadBot with a borrowed platform token rather than any real crawler's string, so it measures how a site responds to the platform token and not how it responds to GPTBot or Googlebot specifically. It also observed no crawler reading anything: no server logs were held, and no figure here says what any answer engine did with either version.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.