BlogFindings

Canonical tag: 163 of 1,069 home pages declared none, and 63 named a URL the page was not served at

Lantad requested the robots.txt and then the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 30 September 2026 as LantadBot, following redirects and executing no JavaScript. 1,069 answered HTTP 200 with an HTML content type. 906 of those declared at least one canonical link element and 163 declared none. Of the 904 that carried a usable href, 841 named exactly the address the page had been served at and 63 named something else, among them a Framer staging host, a different university subdomain and http://localhost:38131/.

20 min read Lantad

This matters for an AI crawler for the same reason it matters for a search engine, and the reason is narrower than it is usually stated. An answer engine that quotes you has to print a link, and the address it prints is the one it decided was yours. Lantad requested the robots.txt and then the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 30 September 2026, sent as its own declared crawler user agent, with redirects followed, a twenty second timeout, from one network location, and no JavaScript executed. 52 hostnames were not asked for a home page at all: 12 served a robots.txt that disallows this scanner at the site root, and a further 40 answered the robots.txt request with a 5xx status, which RFC 9309 treats as an unavailable file that a crawler should read as disallowing everything. 1,069 of the rest answered HTTP 200 with an HTML content type, and what follows counts the canonical link element in those 1,069 responses against the address each response actually came from.

In short

  • Lantad read 1,069 home pages on 30 September 2026 with no JavaScript executed. 906 declared at least one canonical link element, 163 declared none, and 2 of the 906 declared the element with no usable href: affirm.com shipped a canonical link element carrying no href attribute at all, and optainhealth.com shipped one with href set to the empty string.
  • 63 of the 904 usable canonical tags named a URL other than the one the page was served at. 33 differed only in path, 12 only by the www prefix, 6 only in query string, 4 only in scheme, and 7 differed in more than one way, including sj.se, the Swedish national rail operator, whose home page declares its canonical as http://localhost:38131/.
  • 61 home pages answered HTTP 200 with HTML at both the apex and the www address, with identical scheme, path and query, so the same page was reachable at two addresses. 27 of those 61 declared no canonical tag at all, 30 named one of the two, and 4 named a third address that was neither.
  • RFC 6596, published April 2012, defines the canonical link relation as specifying a preferred identifier rather than a binding instruction, and Google's duplicate URL documentation, carrying Last updated 2026-07-10 UTC, calls it a strong signal while stating that none of these methods are required and that Google will pick a version itself if none is given.
  • Presence tracked the platform rather than the sector. All 54 Wix and Squarespace pages and all 37 WordPress small business pages that answered carried a canonical tag, as did 33 of 34 Shopify pages, while 14 of the 42 Webflow pages, 30 of the 103 education pages and 24 of the 85 government pages carried none.
StageCountWhat happened
Hostnames requested1,419The committed corpus, an editorial frame rather than a random draw
robots.txt disallowed this scanner at the site root12Home page not requested for these
robots.txt answered 5xx, read as disallow per RFC 930940Home page not requested for these either
Answered 200 with an HTML content type1,069The denominator for every figure below
Declared at least one canonical link element90684.8 percent of the pages read
Declared none16315.2 percent of the pages read
Declared one with no usable href2affirm.com and optainhealth.com
Named exactly the address served84193.0 percent of the 904 usable
Named a different address637.0 percent of the 904 usable
One GET of https://<host>/robots.txt and then one of https://<host>/ for each hostname, sent as LantadBot/1.0 (+https://lantad.co/bot) with redirects followed and a twenty second timeout, from one network location, with no JavaScript executed. Measured by Lantad on 30 September 2026 across the 1,419 hostnames in worker/seeds/corpus-seeds-industry.json and worker/seeds/corpus-seeds-platform.json.

163 of 1,069 home pages declared no canonical tag

906 of the 1,069 pages carried at least one canonical link element in the head and 163 carried none. Two of the 906 carried the element and nothing usable in it: affirm.com ships a link element with rel set to canonical and no href attribute at all, and optainhealth.com ships one with href set to the empty string. Both were confirmed by a second request. That leaves 904 pages making an actual claim about their own address, and those 904 are the denominator for the next section.

The interesting part is not the headline rate but where the absences sit, because they do not sit where a reader would guess. Absence tracks the platform a site is built on far more closely than it tracks the sector or the size of the organisation. Every one of the 54 Wix and Squarespace pages that answered carried a canonical tag. So did every one of the 37 WordPress small business pages, 33 of the 34 Shopify pages, 30 of the 31 Framer pages and 30 of the 31 static documentation sites. Against that, 14 of the 42 Webflow pages carried none, which makes Webflow the one site builder in this corpus where the tag is frequently absent rather than universally present.

The hand-built end of the corpus is where the absences concentrate. 30 of the 103 education pages, 24 of the 85 government pages and 21 of the 92 healthcare pages carried no canonical tag. Those are the three strata that run the largest and oldest content management systems, and they are the same three that dominated the wrong og:url values in an earlier run. This run did not inspect a single platform's templates and cannot say why any of this is so, but the pattern in the counts is consistent and one reading fits it: the tag is present when something emits it without being asked, and absent when a person has to remember.

It is worth being precise about the cost, because it is easy to overstate. A missing canonical tag is not a defect against the specification, which requires nobody to declare one, and it is not a defect against Google's guidance, which says in its own words that a site will likely do fine without one. What a missing tag does is decline to answer the identity question, leaving the answer to whatever heuristics the consumer runs. That is a reasonable trade when the page has exactly one address. The next section is about the pages where it does not.

StratumPages readDeclared oneDeclared none
wix-squarespace54540
wordpress-smb37370
shopify-dtc34331
framer31301
static-docs31301
news58553
saas1161097
finance937716
healthcare927121
government856124
education1037330
webflow422814
Canonical link element presence on the 1,069 home pages that answered HTTP 200 with an HTML content type, grouped by the corpus stratum each hostname is filed under in the two committed seed files. Strata are editorial, a site is filed by what its pages are built to do or what it is built on, and this is not a random sample of the web. Measured by Lantad on 30 September 2026.

63 canonical tags named a URL the page was not served at

Of the 904 pages carrying a usable canonical, 841 named exactly the address the response had come from, once the fragment was dropped and a trailing slash normalised. 63 named something else. That 7.0 percent is the figure worth carrying away from this post, because every one of those 63 pages is telling a machine that the thing it just fetched is not the authoritative copy.

The 63 divide by how far the declared address sits from the served one. 33 differed only in the path, and these are mostly a content management system exposing its own internals: epa.gov declares /home, box.com declares /home, collibra.com and fasthosts.co.uk each declare /index, and boston.gov declares /homepage-bostongov. 12 differed only by the www prefix, 6 only in the query string, and 4 only in the scheme: ontario.ca, csic.es and bosmun.org each declare an http address on a page served over https, and swapstack.co does the reverse. Those four differ from the served address in transport and nothing else.

Seven differed in more than one way, and this is where the outright errors are. sj.se, the Swedish national rail operator, serves its home page at https://www.sj.se/ and declares its canonical as http://localhost:38131/, an address that exists only on the machine that built the page. blogbowl.io serves at https://www.blogbowl.io/ and declares https://ready-knowledge-301044.framer.app/, a generated staging hostname belonging to the tool the site was built with. mit.edu serves at https://web.mit.edu/ and declares https://tlecms.mit.edu/spotlight/heat-tolerant-vaccines-wednesday, which is not the home page at all but one article on a different subdomain, so the university's front door nominates a single spotlight piece as its own canonical identity. gatech.edu serves at https://www.gatech.edu/ and declares https://gatech.edu/node/1, a raw node path of the kind a content management system generates before a readable alias is applied. All four were confirmed by a second request made after the run.

None of these stops a page being fetched, read or quoted, and it would be an overstatement to say any site here is invisible because of it. The claim is narrower. Each of these pages hands any machine that asks a false answer to the identity question, and the three most striking cases have something in common that is worth naming: a build pipeline, a staging host and a content management system's internal path all leaked into production in a field that renders nothing. The same mechanism produced 340 of 440 sitemap lastmod values that disagreed with the server's own Last-Modified header. Nothing renders it, so nothing checks it.

DifferenceCountExample hostCanonical as returned
Path only33epa.govhttps://www.epa.gov/home
www prefix only12seattle.govhttps://www.seattle.gov/
Query string only6newrelic.comhttps://newrelic.com/?is_us=true
Scheme only4ontario.cahttp://www.ontario.ca/
Host and scheme1sj.sehttp://localhost:38131/
Host only1blogbowl.iohttps://ready-knowledge-301044.framer.app/
Host and path3mit.eduhttps://tlecms.mit.edu/spotlight/heat-tolerant-vaccines-wednesday
Path and query3db.comhttps://www.db.com/index?language_id=1
How the 63 disagreeing canonical values differed from the address the page was served at, with the seven multi-way cases listed in full. Values are quoted exactly as returned. A trailing slash and a fragment were normalised before comparison, so none of these 63 differs only in those. Measured by Lantad on 30 September 2026.

61 home pages answered at two addresses, and 27 named neither as canonical

A canonical tag only earns its place when a page has more than one address, so the second half of this run went looking for pages that do. For each of the 1,069 hostnames that answered, the opposite form of its own hostname was requested once: the www subdomain for a page served at the apex, the apex for a page served at www. 962 of the 1,069 ended at exactly the same address, which is the redirect doing its job and the outcome a site owner should want, and is the opposite of the pattern behind 125 of 1,094 sites answering HTTP 200 for a URL that does not exist. 20 failed to answer and 13 answered with something other than 200 and HTML.

74 answered HTTP 200 with HTML at a second, different address. 61 of those 74 differed from the first address only by the www prefix, with identical scheme, path and query string, and those 61 are the unambiguous cases: one page, two live addresses, neither redirecting to the other. The remaining 13 differed in some other way and several are not duplicates at all but genuinely different pages, admin.ch serving a German home page at one address and an English one at the other, or karolinska.se serving a site selector, so they are excluded from the figures below rather than counted as faults.

Among the 61, the canonical tag is the only thing on the page that can settle which address is real, and on 27 of them there is no canonical tag to do it. boe.es, the Spanish official state gazette, nus.edu.sg, kaist.ac.kr, bb.com.br, bcb.gov.br, turkiye.gov.tr, regions.com and aeromexico.com all serve a working home page at both addresses and declare nothing about which one they prefer. 30 of the 61 do declare a canonical naming one of the two, which is the tag working exactly as intended. The last 4 declare a canonical that names neither address: ama.com.au points at /site/content/default.aspx, maroc.ma points at /ar/node, ug.edu.gh points at /home and spectator.co.uk points at an edition parameter.

Eleven of the 61 sit in the education stratum and eight in healthcare, which is the same constituency as the missing tags in the previous section, and that is not a coincidence so much as the same cause seen twice. A site that never consolidated its two hostnames at the redirect layer is usually a site whose hostnames were configured by different people at different times, which is also the kind of site where a canonical tag was never added. The practical consequence is the one this scanner cares about: a page that is live at two addresses with nothing nominating either is a page whose citations, whatever an engine decides to do with them, can land on two different URLs. That is the same fragmentation problem as 84 of 1,091 home pages offering a crawler no internal path at all, arriving from the opposite direction.

  • Canonical names one of the two 30 of 61 The tag doing the job it exists for
  • No canonical tag at all 27 of 61 Nothing on the page nominates either address
  • Canonical names a third address 4 of 61 ama.com.au, maroc.ma, ug.edu.gh and spectator.co.uk
What the 61 home pages reachable at both their apex and www addresses declare about which address is canonical. The second address was found by requesting the opposite form of each hostname once and comparing the final URL after redirects; all 61 pairs share an identical scheme, path and query string. Measured by Lantad on 30 September 2026.

Three shapes that look wrong and are not

A measurement post is worth less if it counts legal things as faults, and three shapes that a validator or a checklist commonly flags are explicitly permitted by the specification. Counting them here keeps the 63 figure above honest, because none of these three is inside it unless the address also differed.

Relative references are legal. RFC 6596 states the target may specify a relative IRI, and 5 of the 904 usable canonicals were relative: washington.edu and elgiganten.se each declare a single forward slash, fortnox.se declares /home, pnc.com declares /en/personal-banking.html and coalatree.com declares coalatree.com/. Four of those five resolve against the document address to the right place and are correct. Only coalatree.com is wrong, and it is wrong because a bare hostname with no leading slash resolves as a path segment, producing https://coalatree.com/coalatree.com/, which is why it appears in the path-only row above rather than here.

A self-referential canonical is legal and is the overwhelming majority case: 841 of 904. RFC 6596 names it explicitly as permitted, and it is the shape that says the page is the original rather than a copy of something else. A canonical on a different host is also legal, which is the permission that makes the sj.se and blogbowl.io cases interesting rather than merely invalid. Those two are not breaking the specification. They are using a legal construct to point at an address that does not exist for anyone outside the build machine, and no validator that only checks syntax will ever catch it, because there is nothing syntactically wrong.

Six pages declared more than one canonical link element in the head. The specification does not forbid it and does not say what a consumer should do, which leaves the outcome to the consumer's own heuristics, the same escape hatch RFC 6596 grants for improperly declared canonicals generally. Five pages also sent the canonical in an HTTP Link header as well as in the head, and none sent it only there. Anything checking the head alone therefore missed nothing in this corpus, which will not be true of a corpus containing more PDFs and other non-HTML documents, where the header is the only place the relation can go.

  • Self-referential 841 pages, legal RFC 6596 names a self-referential target as permitted. The majority shape and the one that says this page is the original.
  • Relative reference 5 pages, legal Explicitly permitted. Four of the five resolve correctly; coalatree.com omits the leading slash and resolves to a doubled path.
  • Different host 4 pages, legal Permitted by the specification. Syntactically valid in every case here, and pointing at an unreachable address in three of them.
  • More than one element 6 pages, undefined Not forbidden and not specified. RFC 6596 leaves improper declarations to the consumer's own heuristics.
  • Declared in the Link header too 5 pages, legal All five also declared it in the head, so nothing in this corpus was visible only to a consumer reading headers.
  • Element with no usable href 2 pages, broken affirm.com carries no href attribute, optainhealth.com carries an empty one. These declare the relation and say nothing with it.
Shapes found in the 906 declared canonical link elements, against what RFC 6596 permits. Legal means the specification allows the construct, which is a separate question from whether the value is correct. Measured by Lantad on 30 September 2026.

What to check on your own site

Everything in this post was measured from outside with two HTTP requests per hostname and no credentials, which means any of it can be repeated against a site you own in about a minute. The order below is the order the checks actually depend on each other, because a canonical tag is worth checking only after you know how many addresses it has to choose between.

Start with the addresses rather than the tag. Request your home page at the apex and at the www subdomain and follow the redirects on both, then compare the two final URLs. If they match, your redirect layer has already answered the identity question and the canonical tag is a second opinion agreeing with it. If both answer 200 with HTML at different addresses you are one of the 61, and the canonical tag is now the only thing standing between an engine and a guess. Repeat the same comparison for http against https, and for a trailing slash against none.

Then read the tag against the address you were actually served, not against the address you typed. This is the step that catches the failures in this post, because every one of the 63 is a page where the declared value and the served address disagree, and you cannot see that disagreement by looking at the tag alone. The what GPTBot sees view shows the bytes a crawler is handed, and the robots.txt tester covers the access layer that decides whether any of this is reachable in the first place, which the 12 hostnames whose robots.txt disallowed this scanner at the root in this run are a reminder of.

Last, check the value for the three things a syntax validator will not flag: a hostname that only exists inside your build, a staging host belonging to the tool you built with, and a path your content management system uses internally. Those three account for the worst cases here and none of them is malformed. Which address an engine ends up printing is its own decision and an undocumented one, as the notes on getting cited in Google AI Overviews set out. What these three defects have in common with the rest of the machine readable layer, from structured data to the sitemap to the freshness signals an engine may or may not read, is that they are written once by a template and then never looked at by anybody who would recognise a wrong answer. How Lantad weighs each of these layers is set out in the scoring methodology, the corpus-wide figures behind posts like this one sit in the research index, and what any of it is worth for AI visibility depends on an engine behaviour that, on the canonical tag specifically, nobody has yet published.

Sample Illustrative, not a measurement of any real site.

The order of the checks described in this section, as two requests and three comparisons. Every step is performed from outside the site with no credentials. Illustrative of the procedure, not a measurement of any site.

Written by

Lantad

Published .

A canonical tag is one line in the head of a page that answers a question nothing else on the page answers: of all the addresses this content can be reached at, which one is the real one. It is the only part of a normal web page whose whole job is to talk about identity rather than content, and because nothing renders it, nothing breaks when it is wrong. A missing image is noticed within a day. A canonical naming the wrong address can sit in production for years, and this run found several that plainly have.

Common questions

What does a canonical tag do for AI search?

Nothing that any AI operator has published. RFC 6596 defines the canonical link relation as naming a preferred identifier among addresses returning duplicated content, and Google's duplicate URL documentation, Last updated 2026-07-10 UTC, calls it a strong signal while saying it is not required and that Google will choose a version itself if none is declared. No operator documentation behind the crawler tokens this scanner evaluates names the relation either way, so for every engine other than Google the position is undocumented rather than known. This run fetched pages and observed no crawler, so it adds nothing beyond what is in the bytes.

Does every page need a canonical tag?

No specification requires one and Google's own documentation says a site will likely do fine without specifying a preference. 163 of the 1,069 home pages read on 30 September 2026 declared none, and for most of them that is unremarkable because the page has one address. It becomes a real gap when a page is reachable at more than one address: 61 of these home pages answered HTTP 200 with HTML at both their apex and www addresses, and 27 of those 61 declared no canonical tag, so nothing on the page nominates either address.

Can a canonical tag point at a different domain?

Yes. RFC 6596 states the target may exist on a different hostname or domain, so a cross-host canonical is legal rather than a syntax error. That permission is what makes the worst cases in this corpus hard to catch: sj.se declares http://localhost:38131/ and blogbowl.io declares a generated Framer staging hostname, and neither is malformed. A validator checking syntax alone passes both, because the defect is that the address is unreachable from outside the build, not that the value is wrong in shape.

How do I check my own canonical tag?

Request your home page at both the apex and the www subdomain, follow the redirects, and compare the two final URLs. If they differ and both return 200 with HTML, the page has two live addresses and the canonical tag is the only thing that resolves them. Then read the canonical value out of the bytes you were served rather than from the address you typed, because the 63 failures in this corpus are all disagreements between the declared value and the served address. Finally check the value for a build hostname, a staging host or an internal content management system path, which are the three defects no syntax check will flag.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.