BlogFindings
Canonical tag: 163 of 1,069 home pages declared none, and 63 named a URL the page was not served at
Lantad requested the robots.txt and then the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 30 September 2026 as LantadBot, following redirects and executing no JavaScript. 1,069 answered HTTP 200 with an HTML content type. 906 of those declared at least one canonical link element and 163 declared none. Of the 904 that carried a usable href, 841 named exactly the address the page had been served at and 63 named something else, among them a Framer staging host, a different university subdomain and http://localhost:38131/.
This matters for an AI crawler for the same reason it matters for a search engine, and the reason is narrower than it is usually stated. An answer engine that quotes you has to print a link, and the address it prints is the one it decided was yours. Lantad requested the robots.txt and then the home page of all 1,419 hostnames in this repository's two committed corpus seed files on 30 September 2026, sent as its own declared crawler user agent, with redirects followed, a twenty second timeout, from one network location, and no JavaScript executed. 52 hostnames were not asked for a home page at all: 12 served a robots.txt that disallows this scanner at the site root, and a further 40 answered the robots.txt request with a 5xx status, which RFC 9309 treats as an unavailable file that a crawler should read as disallowing everything. 1,069 of the rest answered HTTP 200 with an HTML content type, and what follows counts the canonical link element in those 1,069 responses against the address each response actually came from.
In short
- Lantad read 1,069 home pages on 30 September 2026 with no JavaScript executed. 906 declared at least one canonical link element, 163 declared none, and 2 of the 906 declared the element with no usable href: affirm.com shipped a canonical link element carrying no href attribute at all, and optainhealth.com shipped one with href set to the empty string.
- 63 of the 904 usable canonical tags named a URL other than the one the page was served at. 33 differed only in path, 12 only by the www prefix, 6 only in query string, 4 only in scheme, and 7 differed in more than one way, including sj.se, the Swedish national rail operator, whose home page declares its canonical as http://localhost:38131/.
- 61 home pages answered HTTP 200 with HTML at both the apex and the www address, with identical scheme, path and query, so the same page was reachable at two addresses. 27 of those 61 declared no canonical tag at all, 30 named one of the two, and 4 named a third address that was neither.
- RFC 6596, published April 2012, defines the canonical link relation as specifying a preferred identifier rather than a binding instruction, and Google's duplicate URL documentation, carrying Last updated 2026-07-10 UTC, calls it a strong signal while stating that none of these methods are required and that Google will pick a version itself if none is given.
- Presence tracked the platform rather than the sector. All 54 Wix and Squarespace pages and all 37 WordPress small business pages that answered carried a canonical tag, as did 33 of 34 Shopify pages, while 14 of the 42 Webflow pages, 30 of the 103 education pages and 24 of the 85 government pages carried none.
| Stage | Count | What happened |
|---|---|---|
| Hostnames requested | 1,419 | The committed corpus, an editorial frame rather than a random draw |
| robots.txt disallowed this scanner at the site root | 12 | Home page not requested for these |
| robots.txt answered 5xx, read as disallow per RFC 9309 | 40 | Home page not requested for these either |
| Answered 200 with an HTML content type | 1,069 | The denominator for every figure below |
| Declared at least one canonical link element | 906 | 84.8 percent of the pages read |
| Declared none | 163 | 15.2 percent of the pages read |
| Declared one with no usable href | 2 | affirm.com and optainhealth.com |
| Named exactly the address served | 841 | 93.0 percent of the 904 usable |
| Named a different address | 63 | 7.0 percent of the 904 usable |
What does a canonical tag do for AI search?
The specification is short enough to read in five minutes and it is worth reading before quoting anybody's summary of it. RFC 6596, the canonical link relation, published in April 2012, says the relation specifies the preferred IRI from a set of resources that return the context IRI's content in duplicated form, and it tells authors what to expect from a consumer rather than what a consumer must do. Authors who declare it, the document says, ought to anticipate that applications such as search engines can index content only from the target, consolidate properties such as link popularity to the target, and display the target as the representative identifier. Three permissions in the same section matter for the figures below: the target may be a relative reference, it may be self-referential, and it may exist on a different hostname or domain.
Google is the one operator that publishes a position in enough detail to quote. Google's documentation on consolidating duplicate URLs, carrying Last updated 2026-07-10 UTC, describes a canonical tag as a strong signal that the specified URL should become canonical, then says plainly that none of these methods are required and that a site will likely do fine without specifying a preference, because Google will identify which version to show if the site does not. So on the only published account available, the tag is advisory at both ends: the site is not obliged to send one and the consumer is not obliged to obey it.
The honest frame for everything below follows from that, and it is worth stating before any figure. Nothing in this run shows that an AI crawler read a canonical tag, because no crawler was observed reading anything. This run fetched pages. It did not inspect anybody's retrieval pipeline, it held no server logs, and it made no request from any address other than its own. No operator documentation behind the crawler tokens this scanner evaluates commits to a canonical handling rule at all, which makes the position undocumented rather than settled, and a post that told you otherwise would be inventing the part that is missing.
What can be said without a vendor's help is narrower and still useful. A canonical tag is a claim the page makes about its own identity, which is the same property entity confidence turns on, it is machine readable, it is free to check from outside, and when it is wrong it is wrong in a way nobody on the site will notice. An earlier run of this scanner read canonical tags on five stored captures and found all five self-referential; five pages is too few to say anything about how often that holds, which is the gap this run was built to close. That is the same failure shape as the 14 Open Graph og:url values that did not point at the page that served them, and as the 117 of 612 home pages whose JSON-LD carried a defect: a field bound to the wrong source in a template, shipped, and never read by a human again.
- The relation names a preferred identifier RFC 6596, April 2012. It specifies the preferred IRI from a set of resources returning duplicated content, and tells authors what consumers can do rather than what they must do.
- The target may be on a different host RFC 6596 states the target IRI may exist on a different hostname or domain, and may be relative, and may be self-referential. All three are legal.
- Google calls it a strong signal Google's duplicate URL documentation, Last updated 2026-07-10 UTC. The same page says none of the methods are required and that Google selects a version itself when none is declared.
- An AI crawler honours a canonical tag No operator documentation behind the crawler tokens this scanner evaluates names the canonical link relation either way. Undocumented, not disproved.
- An answer engine cites the canonical address Nothing published settles which address an engine prints when a page is reachable at several. Not observable from outside an engine.
- This run observed a crawler reading a canonical tag It did not. Two GETs per host from one network location, no server logs held, no retrieval pipeline inspected.
163 of 1,069 home pages declared no canonical tag
906 of the 1,069 pages carried at least one canonical link element in the head and 163 carried none. Two of the 906 carried the element and nothing usable in it: affirm.com ships a link element with rel set to canonical and no href attribute at all, and optainhealth.com ships one with href set to the empty string. Both were confirmed by a second request. That leaves 904 pages making an actual claim about their own address, and those 904 are the denominator for the next section.
The interesting part is not the headline rate but where the absences sit, because they do not sit where a reader would guess. Absence tracks the platform a site is built on far more closely than it tracks the sector or the size of the organisation. Every one of the 54 Wix and Squarespace pages that answered carried a canonical tag. So did every one of the 37 WordPress small business pages, 33 of the 34 Shopify pages, 30 of the 31 Framer pages and 30 of the 31 static documentation sites. Against that, 14 of the 42 Webflow pages carried none, which makes Webflow the one site builder in this corpus where the tag is frequently absent rather than universally present.
The hand-built end of the corpus is where the absences concentrate. 30 of the 103 education pages, 24 of the 85 government pages and 21 of the 92 healthcare pages carried no canonical tag. Those are the three strata that run the largest and oldest content management systems, and they are the same three that dominated the wrong og:url values in an earlier run. This run did not inspect a single platform's templates and cannot say why any of this is so, but the pattern in the counts is consistent and one reading fits it: the tag is present when something emits it without being asked, and absent when a person has to remember.
It is worth being precise about the cost, because it is easy to overstate. A missing canonical tag is not a defect against the specification, which requires nobody to declare one, and it is not a defect against Google's guidance, which says in its own words that a site will likely do fine without one. What a missing tag does is decline to answer the identity question, leaving the answer to whatever heuristics the consumer runs. That is a reasonable trade when the page has exactly one address. The next section is about the pages where it does not.
| Stratum | Pages read | Declared one | Declared none |
|---|---|---|---|
| wix-squarespace | 54 | 54 | 0 |
| wordpress-smb | 37 | 37 | 0 |
| shopify-dtc | 34 | 33 | 1 |
| framer | 31 | 30 | 1 |
| static-docs | 31 | 30 | 1 |
| news | 58 | 55 | 3 |
| saas | 116 | 109 | 7 |
| finance | 93 | 77 | 16 |
| healthcare | 92 | 71 | 21 |
| government | 85 | 61 | 24 |
| education | 103 | 73 | 30 |
| webflow | 42 | 28 | 14 |
63 canonical tags named a URL the page was not served at
Of the 904 pages carrying a usable canonical, 841 named exactly the address the response had come from, once the fragment was dropped and a trailing slash normalised. 63 named something else. That 7.0 percent is the figure worth carrying away from this post, because every one of those 63 pages is telling a machine that the thing it just fetched is not the authoritative copy.
The 63 divide by how far the declared address sits from the served one. 33 differed only in the path, and these are mostly a content management system exposing its own internals: epa.gov declares /home, box.com declares /home, collibra.com and fasthosts.co.uk each declare /index, and boston.gov declares /homepage-bostongov. 12 differed only by the www prefix, 6 only in the query string, and 4 only in the scheme: ontario.ca, csic.es and bosmun.org each declare an http address on a page served over https, and swapstack.co does the reverse. Those four differ from the served address in transport and nothing else.
Seven differed in more than one way, and this is where the outright errors are. sj.se, the Swedish national rail operator, serves its home page at https://www.sj.se/ and declares its canonical as http://localhost:38131/, an address that exists only on the machine that built the page. blogbowl.io serves at https://www.blogbowl.io/ and declares https://ready-knowledge-301044.framer.app/, a generated staging hostname belonging to the tool the site was built with. mit.edu serves at https://web.mit.edu/ and declares https://tlecms.mit.edu/spotlight/heat-tolerant-vaccines-wednesday, which is not the home page at all but one article on a different subdomain, so the university's front door nominates a single spotlight piece as its own canonical identity. gatech.edu serves at https://www.gatech.edu/ and declares https://gatech.edu/node/1, a raw node path of the kind a content management system generates before a readable alias is applied. All four were confirmed by a second request made after the run.
None of these stops a page being fetched, read or quoted, and it would be an overstatement to say any site here is invisible because of it. The claim is narrower. Each of these pages hands any machine that asks a false answer to the identity question, and the three most striking cases have something in common that is worth naming: a build pipeline, a staging host and a content management system's internal path all leaked into production in a field that renders nothing. The same mechanism produced 340 of 440 sitemap lastmod values that disagreed with the server's own Last-Modified header. Nothing renders it, so nothing checks it.
| Difference | Count | Example host | Canonical as returned |
|---|---|---|---|
| Path only | 33 | epa.gov | https://www.epa.gov/home |
| www prefix only | 12 | seattle.gov | https://www.seattle.gov/ |
| Query string only | 6 | newrelic.com | https://newrelic.com/?is_us=true |
| Scheme only | 4 | ontario.ca | http://www.ontario.ca/ |
| Host and scheme | 1 | sj.se | http://localhost:38131/ |
| Host only | 1 | blogbowl.io | https://ready-knowledge-301044.framer.app/ |
| Host and path | 3 | mit.edu | https://tlecms.mit.edu/spotlight/heat-tolerant-vaccines-wednesday |
| Path and query | 3 | db.com | https://www.db.com/index?language_id=1 |
61 home pages answered at two addresses, and 27 named neither as canonical
A canonical tag only earns its place when a page has more than one address, so the second half of this run went looking for pages that do. For each of the 1,069 hostnames that answered, the opposite form of its own hostname was requested once: the www subdomain for a page served at the apex, the apex for a page served at www. 962 of the 1,069 ended at exactly the same address, which is the redirect doing its job and the outcome a site owner should want, and is the opposite of the pattern behind 125 of 1,094 sites answering HTTP 200 for a URL that does not exist. 20 failed to answer and 13 answered with something other than 200 and HTML.
74 answered HTTP 200 with HTML at a second, different address. 61 of those 74 differed from the first address only by the www prefix, with identical scheme, path and query string, and those 61 are the unambiguous cases: one page, two live addresses, neither redirecting to the other. The remaining 13 differed in some other way and several are not duplicates at all but genuinely different pages, admin.ch serving a German home page at one address and an English one at the other, or karolinska.se serving a site selector, so they are excluded from the figures below rather than counted as faults.
Among the 61, the canonical tag is the only thing on the page that can settle which address is real, and on 27 of them there is no canonical tag to do it. boe.es, the Spanish official state gazette, nus.edu.sg, kaist.ac.kr, bb.com.br, bcb.gov.br, turkiye.gov.tr, regions.com and aeromexico.com all serve a working home page at both addresses and declare nothing about which one they prefer. 30 of the 61 do declare a canonical naming one of the two, which is the tag working exactly as intended. The last 4 declare a canonical that names neither address: ama.com.au points at /site/content/default.aspx, maroc.ma points at /ar/node, ug.edu.gh points at /home and spectator.co.uk points at an edition parameter.
Eleven of the 61 sit in the education stratum and eight in healthcare, which is the same constituency as the missing tags in the previous section, and that is not a coincidence so much as the same cause seen twice. A site that never consolidated its two hostnames at the redirect layer is usually a site whose hostnames were configured by different people at different times, which is also the kind of site where a canonical tag was never added. The practical consequence is the one this scanner cares about: a page that is live at two addresses with nothing nominating either is a page whose citations, whatever an engine decides to do with them, can land on two different URLs. That is the same fragmentation problem as 84 of 1,091 home pages offering a crawler no internal path at all, arriving from the opposite direction.
Three shapes that look wrong and are not
A measurement post is worth less if it counts legal things as faults, and three shapes that a validator or a checklist commonly flags are explicitly permitted by the specification. Counting them here keeps the 63 figure above honest, because none of these three is inside it unless the address also differed.
Relative references are legal. RFC 6596 states the target may specify a relative IRI, and 5 of the 904 usable canonicals were relative: washington.edu and elgiganten.se each declare a single forward slash, fortnox.se declares /home, pnc.com declares /en/personal-banking.html and coalatree.com declares coalatree.com/. Four of those five resolve against the document address to the right place and are correct. Only coalatree.com is wrong, and it is wrong because a bare hostname with no leading slash resolves as a path segment, producing https://coalatree.com/coalatree.com/, which is why it appears in the path-only row above rather than here.
A self-referential canonical is legal and is the overwhelming majority case: 841 of 904. RFC 6596 names it explicitly as permitted, and it is the shape that says the page is the original rather than a copy of something else. A canonical on a different host is also legal, which is the permission that makes the sj.se and blogbowl.io cases interesting rather than merely invalid. Those two are not breaking the specification. They are using a legal construct to point at an address that does not exist for anyone outside the build machine, and no validator that only checks syntax will ever catch it, because there is nothing syntactically wrong.
Six pages declared more than one canonical link element in the head. The specification does not forbid it and does not say what a consumer should do, which leaves the outcome to the consumer's own heuristics, the same escape hatch RFC 6596 grants for improperly declared canonicals generally. Five pages also sent the canonical in an HTTP Link header as well as in the head, and none sent it only there. Anything checking the head alone therefore missed nothing in this corpus, which will not be true of a corpus containing more PDFs and other non-HTML documents, where the header is the only place the relation can go.
-
Self-referential841 pages, legal RFC 6596 names a self-referential target as permitted. The majority shape and the one that says this page is the original. -
Relative reference5 pages, legal Explicitly permitted. Four of the five resolve correctly; coalatree.com omits the leading slash and resolves to a doubled path. -
Different host4 pages, legal Permitted by the specification. Syntactically valid in every case here, and pointing at an unreachable address in three of them. -
More than one element6 pages, undefined Not forbidden and not specified. RFC 6596 leaves improper declarations to the consumer's own heuristics. -
Declared in the Link header too5 pages, legal All five also declared it in the head, so nothing in this corpus was visible only to a consumer reading headers. -
Element with no usable href2 pages, broken affirm.com carries no href attribute, optainhealth.com carries an empty one. These declare the relation and say nothing with it.
What to check on your own site
Everything in this post was measured from outside with two HTTP requests per hostname and no credentials, which means any of it can be repeated against a site you own in about a minute. The order below is the order the checks actually depend on each other, because a canonical tag is worth checking only after you know how many addresses it has to choose between.
Start with the addresses rather than the tag. Request your home page at the apex and at the www subdomain and follow the redirects on both, then compare the two final URLs. If they match, your redirect layer has already answered the identity question and the canonical tag is a second opinion agreeing with it. If both answer 200 with HTML at different addresses you are one of the 61, and the canonical tag is now the only thing standing between an engine and a guess. Repeat the same comparison for http against https, and for a trailing slash against none.
Then read the tag against the address you were actually served, not against the address you typed. This is the step that catches the failures in this post, because every one of the 63 is a page where the declared value and the served address disagree, and you cannot see that disagreement by looking at the tag alone. The what GPTBot sees view shows the bytes a crawler is handed, and the robots.txt tester covers the access layer that decides whether any of this is reachable in the first place, which the 12 hostnames whose robots.txt disallowed this scanner at the root in this run are a reminder of.
Last, check the value for the three things a syntax validator will not flag: a hostname that only exists inside your build, a staging host belonging to the tool you built with, and a path your content management system uses internally. Those three account for the worst cases here and none of them is malformed. Which address an engine ends up printing is its own decision and an undocumented one, as the notes on getting cited in Google AI Overviews set out. What these three defects have in common with the rest of the machine readable layer, from structured data to the sitemap to the freshness signals an engine may or may not read, is that they are written once by a template and then never looked at by anybody who would recognise a wrong answer. How Lantad weighs each of these layers is set out in the scoring methodology, the corpus-wide figures behind posts like this one sit in the research index, and what any of it is worth for AI visibility depends on an engine behaviour that, on the canonical tag specifically, nobody has yet published.
Sample Illustrative, not a measurement of any real site.
Flow: GET apex home page to Compare final URLs; GET www home page to Compare final URLs; Compare final URLs (same) to One address: redirect settles it; Compare final URLs (differ) to Two live addresses; One address: redirect settles it to Read canonical from served bytes; Two live addresses to Read canonical from served bytes; Read canonical from served bytes to Names the served address; Read canonical from served bytes to Names something else; Names something else to Check for build or staging host.
Lantad
Published .
A canonical tag is one line in the head of a page that answers a question nothing else on the page answers: of all the addresses this content can be reached at, which one is the real one. It is the only part of a normal web page whose whole job is to talk about identity rather than content, and because nothing renders it, nothing breaks when it is wrong. A missing image is noticed within a day. A canonical naming the wrong address can sit in production for years, and this run found several that plainly have.
Common questions
What does a canonical tag do for AI search?
Nothing that any AI operator has published. RFC 6596 defines the canonical link relation as naming a preferred identifier among addresses returning duplicated content, and Google's duplicate URL documentation, Last updated 2026-07-10 UTC, calls it a strong signal while saying it is not required and that Google will choose a version itself if none is declared. No operator documentation behind the crawler tokens this scanner evaluates names the relation either way, so for every engine other than Google the position is undocumented rather than known. This run fetched pages and observed no crawler, so it adds nothing beyond what is in the bytes.
Does every page need a canonical tag?
No specification requires one and Google's own documentation says a site will likely do fine without specifying a preference. 163 of the 1,069 home pages read on 30 September 2026 declared none, and for most of them that is unremarkable because the page has one address. It becomes a real gap when a page is reachable at more than one address: 61 of these home pages answered HTTP 200 with HTML at both their apex and www addresses, and 27 of those 61 declared no canonical tag, so nothing on the page nominates either address.
Can a canonical tag point at a different domain?
Yes. RFC 6596 states the target may exist on a different hostname or domain, so a cross-host canonical is legal rather than a syntax error. That permission is what makes the worst cases in this corpus hard to catch: sj.se declares http://localhost:38131/ and blogbowl.io declares a generated Framer staging hostname, and neither is malformed. A validator checking syntax alone passes both, because the defect is that the address is unreachable from outside the build, not that the value is wrong in shape.
How do I check my own canonical tag?
Request your home page at both the apex and the www subdomain, follow the redirects, and compare the two final URLs. If they differ and both return 200 with HTML, the page has two live addresses and the canonical tag is the only thing that resolves them. Then read the canonical value out of the bytes you were served rather than from the address you typed, because the 63 failures in this corpus are all disagreements between the declared value and the served address. Finally check the value for a build hostname, a staging host or an internal content management system path, which are the three defects no syntax check will flag.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.