BlogFindings

Character encoding for AI search: 15 of 1,088 home pages declared none, and not one of them was a page

The HTML specification requires a character encoding declaration to sit entirely within the first 1024 bytes of a document, because that is all a reader is guaranteed to have examined before it must decide how to turn bytes into text. Lantad read the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 7 October 2026 with no JavaScript executed. 1,088 answered HTTP 200 with HTML, 1,073 of them declared an encoding somewhere, and 1,071 of those declared UTF-8. The 15 that declared nothing were opened one by one, and all 15 turned out to be access denied interstitials, redirect stubs or maintenance notices served with a 200.

19 min read Lantad

Mozilla's reference for the meta element, last modified 24 April 2026 when it was opened for this post, states the requirement plainly: a meta element which declares a character encoding must be located entirely within the first 1024 bytes of the document, and UTF-8 is the only valid encoding for HTML5 documents. The 1024 bytes are not a style preference. They are a budget on how far a reader has to look before it commits to an interpretation of every byte that follows. So this run counted who stays inside it. Lantad read the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 7 October 2026 as LantadBot, following redirects, with no JavaScript executed, and kept the raw bytes rather than a decoded string. 1,088 answered HTTP 200 with an HTML content type. The headline is that this layer is in good order, and the interesting part is what the exceptions turned out to be.

In short

  • Character encoding for AI search is the rare layer of this subject that is already solved: Lantad measured 1,073 of 1,088 readable home pages declaring an encoding on 7 October 2026, and 1,071 of those 1,073 declared UTF-8, with exactly two exceptions in the whole corpus.
  • Mozilla's meta element reference, last modified 24 April 2026, states that a meta element declaring a character encoding must be located entirely within the first 1024 bytes of the document, and that UTF-8 is the only valid encoding for HTML5 documents.
  • 15 of 1,088 readable home pages declared no encoding in the response header or the markup, and opening all 15 found no documents among them: nine were access denied or bot challenge interstitials, four were meta refresh redirect stubs of 92 to 320 bytes, and two were maintenance notices, every one served with HTTP 200.
  • 70 of the 1,034 pages carrying a charset meta element closed that element after byte 1024, at a median of byte 2,334 and as late as byte 40,313, but 62 of the 70 also declared the encoding in the HTTP header, leaving 8 pages whose 231 non-ASCII text characters depend on a reader guessing correctly.
  • jalan.net is the only page in the corpus not served as UTF-8, and the cost of ignoring its declaration is measurable: read as the Windows-31J it declares, its visible text is 7,323 characters with no replacement characters, and read as UTF-8 it is 10,603 characters of which 5,262 are the replacement character.
StageHostsWhat happened
Hostnames asked1,419392 in ten platform strata, 1,027 in eight industry strata
Never returned a status3121 failed in the client, 9 hit the timeout, one terminated
Answered HTTP 403217Refused this crawler
Answered HTTP 50355Served no page to this client
Answered some other status2810 of 202, 7 of 429, 5 of 404, and one each of 400, 401, 406, 451, 498, 500
Answered 200 with HTML1,088The denominator for every rate below
Declared an encoding somewhere1,073906 in the response header, 1,033 in a meta element, 866 in both
Declared UTF-81,071Of the 1,073 that declared anything
Declared nothing anywhere15None of the 15 was a document
Closed the meta after byte 10247062 of them saved by a header declaration
Neither a header nor an in-window meta23The encoding is left to the reader
One GET of https://<host>/ per hostname as LantadBot/1.0 (+https://lantad.co/bot), redirects followed, 20 second timeout, no JavaScript executed, from one network location, with the response body kept as raw bytes. Measured by Lantad on 7 October 2026 across the 1,419 hostnames in this repository's two committed corpus seed files.

How many home pages declare a character encoding?

1,073 of the 1,088 readable home pages declared one somewhere, and 1,071 of those 1,073 declared UTF-8. Two did not. usps.com declared windows-1252 in a meta element at byte 219 with no header declaration at all, and its extracted text contains no non-ASCII characters whatever, so the legacy label costs it nothing today. jalan.net declared Windows-31J in its Content-Type header and shift_jis in its markup, which is the only case in the corpus where a header and a markup declaration name two different encodings, and is not a conflict in practice: both labels resolve to the same decoder, and decoding the page under either produces identical text.

The split across the three declaration sites is worth having, because it is what decides how much the 1024 byte rule actually bites. 906 pages declared the encoding in the Content-Type response header. 1,033 declared a usable value in a charset meta element, and a 1,034th carried the element with an empty value, which is counted separately below. 866 pages declared it in both places, 40 sent a header with no usable meta value, and 167 sent a meta with no header at all. 7 pages opened with a UTF-8 byte order mark, which under the W3C guidance quoted above outranks everything else they went on to say: cbn.gov.ng, federalreserve.gov, nhs.uk, statcan.gc.ca, cancer.ca, sinch.com and schwab.com.

Two forms of the meta element are in use and the proportions are lopsided. 965 pages used the modern charset attribute. 69 used the older http-equiv Content-Type form that Google's special tags page still carries a warning about, and 11 of those 69 also closed the element after byte 1024. 46 pages declared a charset more than once in the same document, up to three times on norebase.com and va.gov, which is harmless under the precedence rules and is a reliable sign of a template assembled from parts rather than written.

What this amounts to is a finding that does not need a fix, and it is worth saying so plainly rather than manufacturing an action from it. The encoding layer of AI visibility is in better shape than the other layers this blog has counted on the same corpus: better than canonical tags, far better than cache validators, and not comparable to the state of structured data. A reader who came here for something to change on their site should check the next two sections and then go and read the crawlability study instead.

  • charset meta element 1034 hosts 965 modern form, 69 legacy http-equiv form
  • Content-Type response header 906 hosts 905 utf-8, one Windows-31J
  • Both header and meta 866 hosts The belt and braces the W3C page recommends
  • Meta only, no header 167 hosts Depends entirely on the 1024 byte window
  • Header only, no meta 40 hosts Safe, and invisible to anyone reading the source
  • UTF-8 byte order mark 7 hosts Outranks every other declaration on the page
  • Nothing anywhere 15 hosts Examined individually in the next section
Where the 1,088 readable home pages declared their character encoding, counted by host. A page can appear in more than one row. Measured by Lantad on 7 October 2026.

What is on the pages that declare no encoding at all?

Fifteen pages declared no encoding in the header and none in the markup. The expectation going in was that these would be old hand written pages, or pages in a language that never needed anything but ASCII. Every one of the fifteen was opened and read, and that is not what they are. Not one of the fifteen is a document.

Nine were refusals. Four of them, statssa.gov.za, pennmedicine.org, heb.com and zurich.com, returned a byte identical 212 byte body carrying a robots meta of noindex,nofollow and a single script element pointing at an Imperva Incapsula resource path, which is a bot challenge wearing a 200. Three more, nus.edu.sg, gtbank.com and regions.com, returned a similar interstitial of 840 to 1,158 bytes with the same noindex,nofollow. lowes.com returned 5,065 bytes whose title element reads Access Denied. nocodemvp.com returned 567 bytes whose title element and only heading both read 403 Forbidden, with nginx named beneath it, under an HTTP status of 200.

Four were redirect stubs that never meant to be read: rbc.com at 148 bytes, irishrail.ie at 92 bytes, contpaqi.com at 210 bytes and sciencespo.fr at 320 bytes, each one a head element holding a meta refresh and, in one case, a noindex. The remaining two, myntra.com and aeromexico.com, served 483 and 551 byte pages whose title element reads Site Maintenance. The largest body in the whole group is lowes.com at 5,065 bytes and the smallest is 92 bytes; the most visible text any of them carried is 16 words and six of them carried none.

This is why the absence of an encoding declaration is worth counting even though the declaration itself is universal. It is not a defect in its own right here. It is a tell. A real template written in the last decade emits the charset line without anyone thinking about it, so a page that lacks the line is usually a page no template produced, which in 2026 means a block page, a challenge or a stub. That makes it a cheap companion signal to the ones this blog has measured directly: the 125 of 1,094 sites that answered a nonexistent path with 200, the five X-Robots-Tag noindex headers that were all a captcha, the 103 of 1,089 sites that served an unknown bot and refused GPTBot, and the GPTBot bans that served a 200 anyway. An HTTP 200 is not evidence that a page was served, and the charset line is one more place that shows.

HostBytesWordsWhat the body actually was
statssa.gov.za2120Incapsula bot challenge, noindex,nofollow
pennmedicine.org2120Incapsula bot challenge, noindex,nofollow
heb.com2120Incapsula bot challenge, noindex,nofollow
zurich.com2120Incapsula bot challenge, noindex,nofollow
nus.edu.sg8526Interstitial, noindex,nofollow
gtbank.com8406Interstitial, noindex,nofollow
regions.com1,1586Interstitial, noindex,nofollow
lowes.com5,06516Title element reads Access Denied
nocodemvp.com5675403 Forbidden body under a 200 status
rbc.com1480Meta refresh stub, noindex
irishrail.ie920Meta refresh stub to /en-ie
contpaqi.com21011Meta refresh stub to the www host
sciencespo.fr3205Meta refresh stub to /fr
myntra.com48310Title element reads Site Maintenance
aeromexico.com55110Title element reads Site Maintenance
Every one of the 15 readable home pages that declared no character encoding, opened and read individually by Lantad on 7 October 2026. Body size is the response in bytes; words is the visible text after scripts, styles and comments are removed.

How many declarations land after the first 1024 bytes?

70 of the 1,034 pages carrying a charset meta element closed that element after byte 1024, which is the one requirement in this subject that sites actually break. The distribution is not a cluster just over the line. The earliest of the 70 closes at byte 1,033 and the median closes at byte 2,334, but 22 of them close after byte 4,096, nine close after byte 10,000, and the latest, temu.com, closes at byte 40,313. A reader that stops looking at 1024 bytes, which is what the requirement implies it may do, finds no declaration on any of the 70.

What pushes the line down the document is consistent once you look. On carecycle.ai the document opens with a decorative banner comment drawn in box drawing characters, which is 1,107 non-ASCII bytes of ornament sitting in front of the declaration that says how to read them. On oliverbonas.com it is the same idea in plain ASCII art, and the declaration does not close until byte 8,320. On dailymaverick.co.za, buildkite.com, scripps.org and webmd.com it is inline script: tag manager snippets, a theme preference check, a bundler's script tags and an advertising preload respectively, all emitted above the charset line. On reliancedigital.in it is the html element itself, which carries an 863 character style attribute of CSS custom properties and so runs to 1,846 bytes before the head element has opened.

62 of the 70 also declared the encoding in the Content-Type header, so the header carries them and the late meta is merely redundant. Eight did not: carecycle.ai, cbn.gov.ng, scripps.org, webmd.com, dailymaverick.co.za, buildkite.com, oliverbonas.com and reliancedigital.in. Those eight, plus the fifteen that declared nothing, are the 23 pages of 1,088 where a reader honouring the window has nothing to go on.

The eight are not a disaster and the honest figure is small. Their extracted text holds 231 non-ASCII characters between them, and the heaviest is dailymaverick.co.za with 104. What those characters are is the part worth knowing, because it contradicts the intuition that this is a problem for sites in other languages. They are punctuation. 56 of dailymaverick.co.za's 104 are a right single quotation mark doing duty as an apostrophe, and the rest run to em dashes, quotation marks and three accented capitals. 34 of oliverbonas.com's 37 are a pound sign. buildkite.com's 26 are mostly arrows and curly quotes. The pages at risk here are English language pages, and what they risk is their apostrophes and their prices, which is also why this rarely gets noticed. The one stratum that stands out is Shopify, where 8 of 34 readable storefronts closed the declaration late, a pattern that belongs with the other template level findings in the Shopify fix guide and in Shopify AI crawlers.

HostMeta closes at byteNon-ASCII text charactersWhat sits in front of the declaration
dailymaverick.co.za1,917104Inline tag manager snippets
oliverbonas.com8,32037An ASCII art banner comment
carecycle.ai1,39632A banner comment in box drawing characters
buildkite.com1,75826An inline theme preference script
scripps.org7,76226Inline script and bundler script tags
webmd.com1,3106An advertising preload and inline script
cbn.gov.ng1,8100Opens with a byte order mark, which outranks it
reliancedigital.in1,9060CSS custom properties in a style attribute
The 8 home pages that closed their charset meta element after byte 1024 and sent no charset in the Content-Type header, so a reader honouring the 1024 byte window has no declaration to use. Non-ASCII text characters are counted in the visible text after the document is decoded as UTF-8. Measured by Lantad on 7 October 2026.

What does ignoring the declaration actually cost?

One page in the corpus lets this be priced rather than argued. jalan.net, a Japanese travel site, is the only one of the 1,088 that is not served as UTF-8. Its Content-Type header declares Windows-31J and its markup declares shift_jis, and its bytes are correspondingly not valid UTF-8 at all.

Read as it asks to be read, the page yields 7,323 characters of visible text, 1,200 words, and not one replacement character. Its title resolves to the Japanese for accommodation and hotel booking followed by the site's name. Read as UTF-8, the same bytes yield 10,603 characters of which 5,262 are U+FFFD, the replacement character, and the title resolves to nothing legible at all. That gap, 5,262 characters on one page, is the entire value of honouring an encoding declaration, stated as a number. It is also the reason the corpus wide figures matter: across all 1,088 readable home pages the extracted text runs to 9,730,690 characters, of which 156,488 are non-ASCII and 118,953 sit beyond the Latin-1 range, and 922 of the 1,088 pages carry at least one character in that upper range. Those are the characters whose identity depends on getting the encoding right.

There is one page where the damage has already happened upstream and is now permanent. absa.co.za declares UTF-8, serves valid UTF-8, and carries exactly one U+FFFD inside its own bytes, in a span whose class attribute is copyright-symbol, where a copyright sign should be. Nothing a reader does can recover that character, because the page is faithfully serving a replacement character that was baked in when some earlier tool decoded the text wrongly and re-encoded the result. It is a single character on a single page and it is the clearest illustration in the corpus of what this failure mode looks like after the fact: not a broken page, just a word that is quietly no longer the word.

Two things this measurement deliberately does not claim. It does not claim that any AI crawler mishandles jalan.net, because no crawler was observed on it; what was measured is what the bytes contain and what two readings of them produce. And it does not claim that encoding problems are a meaningful share of lost text on the web. On this corpus they are not, and the volumes that actually go missing are elsewhere: the 121,312 of 1,269,054 words sitting behind a hiding marker, the 1,665 of 13,615 words lost by deleting one main element, and the 17 of 380 home pages that sent a crawler no words at all.

Decoded as declared

  • Content-Type: text/html;charset=Windows-31J
  • 7,323 characters of visible text
  • 1,200 words
  • 0 replacement characters
  • Title resolves to readable Japanese

Decoded as UTF-8

  • Declaration ignored
  • 10,603 characters of visible text
  • Word count not meaningful
  • 5,262 replacement characters
  • Title resolves to nothing legible
jalan.net, the only home page of the 1,088 not served as UTF-8, read twice from the same bytes captured by Lantad on 7 October 2026: once under the Windows-31J it declares and once under UTF-8.

The declarations that are present and still wrong

A charset line can be in the right place and still say nothing useful, and two pages in the corpus show the two ways that happens. patsnap.com emits a meta element reading charset="Patsnap", the company's own name in the slot reserved for an encoding label, at byte 3,828. motherjones.com emits a meta element with an empty charset value, charset="", at byte 190. Neither value names an encoding. Both pages are rescued by their Content-Type header, which declares UTF-8 in each case, and that is the whole argument for the belt and braces approach the W3C page recommends: 866 pages in this corpus declare the encoding twice, and on these two the second declaration is the only one that works.

Neither of these is detectable by looking at a page in a browser, which is the recurring difficulty with this entire class of defect. A browser that has a usable declaration in the header never reveals that the one in the markup is junk. The same is true of the late declarations in the previous section and of the duplicate declarations: everything renders, every check that uses a browser passes, and the only way to see any of it is to read the raw bytes in the order a reader reads them. That is the same reason a browser based check misses the findings in what a crawler sees without JavaScript and why the methodology page is explicit about fetching bytes rather than rendering pages.

It is worth separating this from the questions about language that it gets confused with. An encoding declaration says how to turn bytes into characters. It says nothing about which language those characters are in, which is a separate declaration that 54 of the 1,088 pages omitted entirely, and nothing about serving different languages to different readers, which this blog measured in hreflang and AI crawlers and in Googlebot setting no Accept-Language. It is also separate from naming the language inside structured data, where 429 of 615 pages carrying parseable JSON-LD named no language at all. Encoding is the layer below all of those and it is the one that is working.

So the practical reading of this run is short. Declare UTF-8 in the Content-Type header as well as in the markup, and the 1024 byte rule stops being able to hurt you. Put the meta charset element immediately after the opening head tag, before any script, any comment and any style attribute, which costs nothing and is what the W3C guidance asks for. If you want to know whether your own site does either of those things, the request Lantad makes is documented at the bot page, and the single GET it takes to check is the same one behind every figure above.

  • charset="Patsnap" 1 host patsnap.com names itself where an encoding label belongs, at byte 3,828. The UTF-8 in its header is what saves it.
  • charset="" 1 host motherjones.com declares an empty value at byte 190. Its header declares UTF-8 and carries the page.
  • Declared more than once 46 hosts Harmless under the precedence rules, and a reliable sign of a template assembled from parts. Three times on norebase.com and va.gov.
  • Legacy http-equiv form 69 hosts The form Google's special tags page still warns about quoting correctly. 11 of the 69 also closed after byte 1024.
The four malformed or redundant declaration patterns found across the 1,088 readable home pages, with the host count for each. Measured by Lantad on 7 October 2026.

Written by

Lantad

Published .

Almost every question about whether an AI crawler can read a site turns out to be a question about access or about rendering: whether the request was answered, and whether the words arrived without a browser. Character encoding for AI search sits underneath both of those, and it is usually skipped, because it looks like a problem the web solved fifteen years ago. One line in the head, UTF-8, done. The reason it is worth measuring anyway is that the rule attached to that line is stricter than most people remember, and the pages that break it are not random.

Common questions

Does character encoding affect AI search visibility?

Only in one narrow way, and on this corpus it is close to settled. An encoding declaration decides what the bytes mean once they arrive, not whether a crawler may fetch the page or whether the words are in the HTML. Lantad measured 1,073 of 1,088 readable home pages declaring an encoding on 7 October 2026, and 1,071 of those declared UTF-8, so for almost every site the answer is that this layer is already correct and the text that goes missing goes missing somewhere else.

Where should the meta charset tag go?

Immediately after the opening head tag, before any script, comment or style attribute. Mozilla's meta element reference, last modified 24 April 2026, states that a meta element declaring a character encoding must be located entirely within the first 1024 bytes of the document, and the W3C internationalisation guidance recommends that position for exactly that reason. Lantad found 70 of 1,034 pages closing the element after byte 1024 on 7 October 2026, at a median of byte 2,334 and as late as byte 40,313.

Is declaring the encoding in the HTTP header enough on its own?

It is sufficient under the precedence rules, because the W3C internationalisation guidance states that the HTTP header has the highest priority when it conflicts with in-document declarations other than a byte order mark. In practice declaring it in both places is what protects you, and this run shows why: patsnap.com and motherjones.com both emit a meta charset element that names no encoding, and both are readable only because their Content-Type header declares UTF-8.

What happens if a page declares no character encoding at all?

The reader has to decide for itself, and the documentation read for this post does not specify what it should decide. Lantad found 15 of 1,088 readable home pages declaring nothing anywhere on 7 October 2026, and opening all 15 found no documents among them: nine access denied or bot challenge interstitials, four meta refresh stubs and two maintenance notices, all served with HTTP 200. A missing charset line is better read as a signal that no real page was served than as an encoding defect.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.