BlogFindings
Character encoding for AI search: 15 of 1,088 home pages declared none, and not one of them was a page
The HTML specification requires a character encoding declaration to sit entirely within the first 1024 bytes of a document, because that is all a reader is guaranteed to have examined before it must decide how to turn bytes into text. Lantad read the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 7 October 2026 with no JavaScript executed. 1,088 answered HTTP 200 with HTML, 1,073 of them declared an encoding somewhere, and 1,071 of those declared UTF-8. The 15 that declared nothing were opened one by one, and all 15 turned out to be access denied interstitials, redirect stubs or maintenance notices served with a 200.
Mozilla's reference for the meta element, last modified 24 April 2026 when it was opened for this post, states the requirement plainly: a meta element which declares a character encoding must be located entirely within the first 1024 bytes of the document, and UTF-8 is the only valid encoding for HTML5 documents. The 1024 bytes are not a style preference. They are a budget on how far a reader has to look before it commits to an interpretation of every byte that follows. So this run counted who stays inside it. Lantad read the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 7 October 2026 as LantadBot, following redirects, with no JavaScript executed, and kept the raw bytes rather than a decoded string. 1,088 answered HTTP 200 with an HTML content type. The headline is that this layer is in good order, and the interesting part is what the exceptions turned out to be.
In short
- Character encoding for AI search is the rare layer of this subject that is already solved: Lantad measured 1,073 of 1,088 readable home pages declaring an encoding on 7 October 2026, and 1,071 of those 1,073 declared UTF-8, with exactly two exceptions in the whole corpus.
- Mozilla's meta element reference, last modified 24 April 2026, states that a meta element declaring a character encoding must be located entirely within the first 1024 bytes of the document, and that UTF-8 is the only valid encoding for HTML5 documents.
- 15 of 1,088 readable home pages declared no encoding in the response header or the markup, and opening all 15 found no documents among them: nine were access denied or bot challenge interstitials, four were meta refresh redirect stubs of 92 to 320 bytes, and two were maintenance notices, every one served with HTTP 200.
- 70 of the 1,034 pages carrying a charset meta element closed that element after byte 1024, at a median of byte 2,334 and as late as byte 40,313, but 62 of the 70 also declared the encoding in the HTTP header, leaving 8 pages whose 231 non-ASCII text characters depend on a reader guessing correctly.
- jalan.net is the only page in the corpus not served as UTF-8, and the cost of ignoring its declaration is measurable: read as the Windows-31J it declares, its visible text is 7,323 characters with no replacement characters, and read as UTF-8 it is 10,603 characters of which 5,262 are the replacement character.
| Stage | Hosts | What happened |
|---|---|---|
| Hostnames asked | 1,419 | 392 in ten platform strata, 1,027 in eight industry strata |
| Never returned a status | 31 | 21 failed in the client, 9 hit the timeout, one terminated |
| Answered HTTP 403 | 217 | Refused this crawler |
| Answered HTTP 503 | 55 | Served no page to this client |
| Answered some other status | 28 | 10 of 202, 7 of 429, 5 of 404, and one each of 400, 401, 406, 451, 498, 500 |
| Answered 200 with HTML | 1,088 | The denominator for every rate below |
| Declared an encoding somewhere | 1,073 | 906 in the response header, 1,033 in a meta element, 866 in both |
| Declared UTF-8 | 1,071 | Of the 1,073 that declared anything |
| Declared nothing anywhere | 15 | None of the 15 was a document |
| Closed the meta after byte 1024 | 70 | 62 of them saved by a header declaration |
| Neither a header nor an in-window meta | 23 | The encoding is left to the reader |
Does character encoding matter for AI search?
It matters in exactly one way, and the way is worth stating precisely, because it is narrower than the usual advice implies. An encoding declaration does not affect whether a crawler is allowed to fetch the page, which is settled earlier by the layers described in two layers decide if AI can read your site. It does not affect whether the words are in the HTML at all, which is the prose parity question. It decides only what the bytes mean once they have arrived, and it therefore decides whether the text a reader extracts is the text an author wrote.
Three places can carry the declaration, and they are not equal. The W3C's internationalisation guidance on encoding declarations states that the HTTP header information has the highest priority when it conflicts with in-document declarations other than the byte order mark, and that a UTF-8 byte order mark at the start of a file has a higher precedence than any other declaration, including the HTTP header. The same page says the in-document declaration should fit completely within the first 1024 bytes at the start of the file, and recommends putting it immediately after the opening head tag. So the order is byte order mark, then header, then markup, and the markup form is the only one with a position requirement, because it is the only one a reader has to go looking for.
Google's own guidance touches this too, from a different angle. Google's page on special tags, Last updated 2025-12-10 UTC, recommends Unicode and UTF-8 where possible, and warns that the value of the content attribute in the http-equiv form of the tag has to be surrounded by quotes or the charset may be interpreted incorrectly. That is a note about a legacy syntax rather than a ranking factor, and it is in the documentation because the legacy syntax is still in use. 69 pages in this corpus still use it.
None of this is a measurement of crawler behaviour, and this post does not make one. No crawler was observed decoding anything here. What was measured is what 1,088 sites put on the wire, read against what the documentation above requires of them, which is the same division this blog keeps when it reports what a vendor says against what sites do in meta tags for AI search.
Flow: Response bytes arrive (checked first) to UTF-8 byte order mark present; Response bytes arrive (no mark) to charset in Content-Type header; Response bytes arrive (no header) to charset meta inside first 1024 bytes; Response bytes arrive (none of the three) to Nothing declared, reader decides; UTF-8 byte order mark present to Bytes become text; charset in Content-Type header to Bytes become text; charset meta inside first 1024 bytes to Bytes become text; Nothing declared, reader decides (may differ) to Bytes become text.
How many home pages declare a character encoding?
1,073 of the 1,088 readable home pages declared one somewhere, and 1,071 of those 1,073 declared UTF-8. Two did not. usps.com declared windows-1252 in a meta element at byte 219 with no header declaration at all, and its extracted text contains no non-ASCII characters whatever, so the legacy label costs it nothing today. jalan.net declared Windows-31J in its Content-Type header and shift_jis in its markup, which is the only case in the corpus where a header and a markup declaration name two different encodings, and is not a conflict in practice: both labels resolve to the same decoder, and decoding the page under either produces identical text.
The split across the three declaration sites is worth having, because it is what decides how much the 1024 byte rule actually bites. 906 pages declared the encoding in the Content-Type response header. 1,033 declared a usable value in a charset meta element, and a 1,034th carried the element with an empty value, which is counted separately below. 866 pages declared it in both places, 40 sent a header with no usable meta value, and 167 sent a meta with no header at all. 7 pages opened with a UTF-8 byte order mark, which under the W3C guidance quoted above outranks everything else they went on to say: cbn.gov.ng, federalreserve.gov, nhs.uk, statcan.gc.ca, cancer.ca, sinch.com and schwab.com.
Two forms of the meta element are in use and the proportions are lopsided. 965 pages used the modern charset attribute. 69 used the older http-equiv Content-Type form that Google's special tags page still carries a warning about, and 11 of those 69 also closed the element after byte 1024. 46 pages declared a charset more than once in the same document, up to three times on norebase.com and va.gov, which is harmless under the precedence rules and is a reliable sign of a template assembled from parts rather than written.
What this amounts to is a finding that does not need a fix, and it is worth saying so plainly rather than manufacturing an action from it. The encoding layer of AI visibility is in better shape than the other layers this blog has counted on the same corpus: better than canonical tags, far better than cache validators, and not comparable to the state of structured data. A reader who came here for something to change on their site should check the next two sections and then go and read the crawlability study instead.
What is on the pages that declare no encoding at all?
Fifteen pages declared no encoding in the header and none in the markup. The expectation going in was that these would be old hand written pages, or pages in a language that never needed anything but ASCII. Every one of the fifteen was opened and read, and that is not what they are. Not one of the fifteen is a document.
Nine were refusals. Four of them, statssa.gov.za, pennmedicine.org, heb.com and zurich.com, returned a byte identical 212 byte body carrying a robots meta of noindex,nofollow and a single script element pointing at an Imperva Incapsula resource path, which is a bot challenge wearing a 200. Three more, nus.edu.sg, gtbank.com and regions.com, returned a similar interstitial of 840 to 1,158 bytes with the same noindex,nofollow. lowes.com returned 5,065 bytes whose title element reads Access Denied. nocodemvp.com returned 567 bytes whose title element and only heading both read 403 Forbidden, with nginx named beneath it, under an HTTP status of 200.
Four were redirect stubs that never meant to be read: rbc.com at 148 bytes, irishrail.ie at 92 bytes, contpaqi.com at 210 bytes and sciencespo.fr at 320 bytes, each one a head element holding a meta refresh and, in one case, a noindex. The remaining two, myntra.com and aeromexico.com, served 483 and 551 byte pages whose title element reads Site Maintenance. The largest body in the whole group is lowes.com at 5,065 bytes and the smallest is 92 bytes; the most visible text any of them carried is 16 words and six of them carried none.
This is why the absence of an encoding declaration is worth counting even though the declaration itself is universal. It is not a defect in its own right here. It is a tell. A real template written in the last decade emits the charset line without anyone thinking about it, so a page that lacks the line is usually a page no template produced, which in 2026 means a block page, a challenge or a stub. That makes it a cheap companion signal to the ones this blog has measured directly: the 125 of 1,094 sites that answered a nonexistent path with 200, the five X-Robots-Tag noindex headers that were all a captcha, the 103 of 1,089 sites that served an unknown bot and refused GPTBot, and the GPTBot bans that served a 200 anyway. An HTTP 200 is not evidence that a page was served, and the charset line is one more place that shows.
| Host | Bytes | Words | What the body actually was |
|---|---|---|---|
| statssa.gov.za | 212 | 0 | Incapsula bot challenge, noindex,nofollow |
| pennmedicine.org | 212 | 0 | Incapsula bot challenge, noindex,nofollow |
| heb.com | 212 | 0 | Incapsula bot challenge, noindex,nofollow |
| zurich.com | 212 | 0 | Incapsula bot challenge, noindex,nofollow |
| nus.edu.sg | 852 | 6 | Interstitial, noindex,nofollow |
| gtbank.com | 840 | 6 | Interstitial, noindex,nofollow |
| regions.com | 1,158 | 6 | Interstitial, noindex,nofollow |
| lowes.com | 5,065 | 16 | Title element reads Access Denied |
| nocodemvp.com | 567 | 5 | 403 Forbidden body under a 200 status |
| rbc.com | 148 | 0 | Meta refresh stub, noindex |
| irishrail.ie | 92 | 0 | Meta refresh stub to /en-ie |
| contpaqi.com | 210 | 11 | Meta refresh stub to the www host |
| sciencespo.fr | 320 | 5 | Meta refresh stub to /fr |
| myntra.com | 483 | 10 | Title element reads Site Maintenance |
| aeromexico.com | 551 | 10 | Title element reads Site Maintenance |
How many declarations land after the first 1024 bytes?
70 of the 1,034 pages carrying a charset meta element closed that element after byte 1024, which is the one requirement in this subject that sites actually break. The distribution is not a cluster just over the line. The earliest of the 70 closes at byte 1,033 and the median closes at byte 2,334, but 22 of them close after byte 4,096, nine close after byte 10,000, and the latest, temu.com, closes at byte 40,313. A reader that stops looking at 1024 bytes, which is what the requirement implies it may do, finds no declaration on any of the 70.
What pushes the line down the document is consistent once you look. On carecycle.ai the document opens with a decorative banner comment drawn in box drawing characters, which is 1,107 non-ASCII bytes of ornament sitting in front of the declaration that says how to read them. On oliverbonas.com it is the same idea in plain ASCII art, and the declaration does not close until byte 8,320. On dailymaverick.co.za, buildkite.com, scripps.org and webmd.com it is inline script: tag manager snippets, a theme preference check, a bundler's script tags and an advertising preload respectively, all emitted above the charset line. On reliancedigital.in it is the html element itself, which carries an 863 character style attribute of CSS custom properties and so runs to 1,846 bytes before the head element has opened.
62 of the 70 also declared the encoding in the Content-Type header, so the header carries them and the late meta is merely redundant. Eight did not: carecycle.ai, cbn.gov.ng, scripps.org, webmd.com, dailymaverick.co.za, buildkite.com, oliverbonas.com and reliancedigital.in. Those eight, plus the fifteen that declared nothing, are the 23 pages of 1,088 where a reader honouring the window has nothing to go on.
The eight are not a disaster and the honest figure is small. Their extracted text holds 231 non-ASCII characters between them, and the heaviest is dailymaverick.co.za with 104. What those characters are is the part worth knowing, because it contradicts the intuition that this is a problem for sites in other languages. They are punctuation. 56 of dailymaverick.co.za's 104 are a right single quotation mark doing duty as an apostrophe, and the rest run to em dashes, quotation marks and three accented capitals. 34 of oliverbonas.com's 37 are a pound sign. buildkite.com's 26 are mostly arrows and curly quotes. The pages at risk here are English language pages, and what they risk is their apostrophes and their prices, which is also why this rarely gets noticed. The one stratum that stands out is Shopify, where 8 of 34 readable storefronts closed the declaration late, a pattern that belongs with the other template level findings in the Shopify fix guide and in Shopify AI crawlers.
| Host | Meta closes at byte | Non-ASCII text characters | What sits in front of the declaration |
|---|---|---|---|
| dailymaverick.co.za | 1,917 | 104 | Inline tag manager snippets |
| oliverbonas.com | 8,320 | 37 | An ASCII art banner comment |
| carecycle.ai | 1,396 | 32 | A banner comment in box drawing characters |
| buildkite.com | 1,758 | 26 | An inline theme preference script |
| scripps.org | 7,762 | 26 | Inline script and bundler script tags |
| webmd.com | 1,310 | 6 | An advertising preload and inline script |
| cbn.gov.ng | 1,810 | 0 | Opens with a byte order mark, which outranks it |
| reliancedigital.in | 1,906 | 0 | CSS custom properties in a style attribute |
What does ignoring the declaration actually cost?
One page in the corpus lets this be priced rather than argued. jalan.net, a Japanese travel site, is the only one of the 1,088 that is not served as UTF-8. Its Content-Type header declares Windows-31J and its markup declares shift_jis, and its bytes are correspondingly not valid UTF-8 at all.
Read as it asks to be read, the page yields 7,323 characters of visible text, 1,200 words, and not one replacement character. Its title resolves to the Japanese for accommodation and hotel booking followed by the site's name. Read as UTF-8, the same bytes yield 10,603 characters of which 5,262 are U+FFFD, the replacement character, and the title resolves to nothing legible at all. That gap, 5,262 characters on one page, is the entire value of honouring an encoding declaration, stated as a number. It is also the reason the corpus wide figures matter: across all 1,088 readable home pages the extracted text runs to 9,730,690 characters, of which 156,488 are non-ASCII and 118,953 sit beyond the Latin-1 range, and 922 of the 1,088 pages carry at least one character in that upper range. Those are the characters whose identity depends on getting the encoding right.
There is one page where the damage has already happened upstream and is now permanent. absa.co.za declares UTF-8, serves valid UTF-8, and carries exactly one U+FFFD inside its own bytes, in a span whose class attribute is copyright-symbol, where a copyright sign should be. Nothing a reader does can recover that character, because the page is faithfully serving a replacement character that was baked in when some earlier tool decoded the text wrongly and re-encoded the result. It is a single character on a single page and it is the clearest illustration in the corpus of what this failure mode looks like after the fact: not a broken page, just a word that is quietly no longer the word.
Two things this measurement deliberately does not claim. It does not claim that any AI crawler mishandles jalan.net, because no crawler was observed on it; what was measured is what the bytes contain and what two readings of them produce. And it does not claim that encoding problems are a meaningful share of lost text on the web. On this corpus they are not, and the volumes that actually go missing are elsewhere: the 121,312 of 1,269,054 words sitting behind a hiding marker, the 1,665 of 13,615 words lost by deleting one main element, and the 17 of 380 home pages that sent a crawler no words at all.
Decoded as declared
- Content-Type: text/html;charset=Windows-31J
- 7,323 characters of visible text
- 1,200 words
- 0 replacement characters
- Title resolves to readable Japanese
Decoded as UTF-8
- Declaration ignored
- 10,603 characters of visible text
- Word count not meaningful
- 5,262 replacement characters
- Title resolves to nothing legible
The declarations that are present and still wrong
A charset line can be in the right place and still say nothing useful, and two pages in the corpus show the two ways that happens. patsnap.com emits a meta element reading charset="Patsnap", the company's own name in the slot reserved for an encoding label, at byte 3,828. motherjones.com emits a meta element with an empty charset value, charset="", at byte 190. Neither value names an encoding. Both pages are rescued by their Content-Type header, which declares UTF-8 in each case, and that is the whole argument for the belt and braces approach the W3C page recommends: 866 pages in this corpus declare the encoding twice, and on these two the second declaration is the only one that works.
Neither of these is detectable by looking at a page in a browser, which is the recurring difficulty with this entire class of defect. A browser that has a usable declaration in the header never reveals that the one in the markup is junk. The same is true of the late declarations in the previous section and of the duplicate declarations: everything renders, every check that uses a browser passes, and the only way to see any of it is to read the raw bytes in the order a reader reads them. That is the same reason a browser based check misses the findings in what a crawler sees without JavaScript and why the methodology page is explicit about fetching bytes rather than rendering pages.
It is worth separating this from the questions about language that it gets confused with. An encoding declaration says how to turn bytes into characters. It says nothing about which language those characters are in, which is a separate declaration that 54 of the 1,088 pages omitted entirely, and nothing about serving different languages to different readers, which this blog measured in hreflang and AI crawlers and in Googlebot setting no Accept-Language. It is also separate from naming the language inside structured data, where 429 of 615 pages carrying parseable JSON-LD named no language at all. Encoding is the layer below all of those and it is the one that is working.
So the practical reading of this run is short. Declare UTF-8 in the Content-Type header as well as in the markup, and the 1024 byte rule stops being able to hurt you. Put the meta charset element immediately after the opening head tag, before any script, any comment and any style attribute, which costs nothing and is what the W3C guidance asks for. If you want to know whether your own site does either of those things, the request Lantad makes is documented at the bot page, and the single GET it takes to check is the same one behind every figure above.
-
charset="Patsnap"1 host patsnap.com names itself where an encoding label belongs, at byte 3,828. The UTF-8 in its header is what saves it. -
charset=""1 host motherjones.com declares an empty value at byte 190. Its header declares UTF-8 and carries the page. -
Declared more than once46 hosts Harmless under the precedence rules, and a reliable sign of a template assembled from parts. Three times on norebase.com and va.gov. -
Legacy http-equiv form69 hosts The form Google's special tags page still warns about quoting correctly. 11 of the 69 also closed after byte 1024.
Lantad
Published .
Almost every question about whether an AI crawler can read a site turns out to be a question about access or about rendering: whether the request was answered, and whether the words arrived without a browser. Character encoding for AI search sits underneath both of those, and it is usually skipped, because it looks like a problem the web solved fifteen years ago. One line in the head, UTF-8, done. The reason it is worth measuring anyway is that the rule attached to that line is stricter than most people remember, and the pages that break it are not random.
Common questions
Does character encoding affect AI search visibility?
Only in one narrow way, and on this corpus it is close to settled. An encoding declaration decides what the bytes mean once they arrive, not whether a crawler may fetch the page or whether the words are in the HTML. Lantad measured 1,073 of 1,088 readable home pages declaring an encoding on 7 October 2026, and 1,071 of those declared UTF-8, so for almost every site the answer is that this layer is already correct and the text that goes missing goes missing somewhere else.
Where should the meta charset tag go?
Immediately after the opening head tag, before any script, comment or style attribute. Mozilla's meta element reference, last modified 24 April 2026, states that a meta element declaring a character encoding must be located entirely within the first 1024 bytes of the document, and the W3C internationalisation guidance recommends that position for exactly that reason. Lantad found 70 of 1,034 pages closing the element after byte 1024 on 7 October 2026, at a median of byte 2,334 and as late as byte 40,313.
Is declaring the encoding in the HTTP header enough on its own?
It is sufficient under the precedence rules, because the W3C internationalisation guidance states that the HTTP header has the highest priority when it conflicts with in-document declarations other than a byte order mark. In practice declaring it in both places is what protects you, and this run shows why: patsnap.com and motherjones.com both emit a meta charset element that names no encoding, and both are readable only because their Content-Type header declares UTF-8.
What happens if a page declares no character encoding at all?
The reader has to decide for itself, and the documentation read for this post does not specify what it should decide. Lantad found 15 of 1,088 readable home pages declaring nothing anywhere on 7 October 2026, and opening all 15 found no documents among them: nine access denied or bot challenge interstitials, four meta refresh stubs and two maintenance notices, all served with HTTP 200. A missing charset line is better read as a signal that no real page was served than as an encoding defect.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.