BlogFindings
Do AI crawlers read SVG text? 80 of 37,019 inline graphics carried any
Lantad asked all 1,419 hostnames in this repository's two committed corpus seed files for robots.txt on 6 October 2026 and read every permitted home page as LantadBot with no JavaScript executed. 1,074 answered HTTP 200 with HTML and carried 37,019 inline SVG elements between them. 80 of those elements carried a text, tspan or textPath node, the words inside them came to 1,285 against 1,355,047 words of ordinary page text, and 35,263 of the 37,019 carried no text of any kind.
So this run counted what the line costs. All 1,419 hostnames in the two committed corpus seed files were asked for robots.txt and then for their home page, once each, as LantadBot, on 6 October 2026, with no JavaScript executed and no stylesheet fetched. 1,074 answered HTTP 200 with HTML. Every inline SVG element in those pages was located with nesting awareness, and each one was read twice: once for the characters inside its text, tspan and textPath nodes, which are the parts a person sees, and once for its title and desc children, which are the parts only assistive technology and a careful parser see. The result is lopsided enough to be useful in both directions, because the aggregate vindicates the decision and a short list of pages does not. It also lands on this site's own markup, since the diagrams in this post are inline SVG and the AI crawler advice on prose parity is written by the same people who wrote the drop rule. The nearest prior run, on table elements, asked a supply-side question and got a supply-side answer. This one asks what the scanner throws away.
In short
- Do AI crawlers read SVG text is a question about 1,285 words: of the 37,019 inline SVG elements Lantad read across 1,074 corpus home pages on 6 October 2026, 80 carried a text, tspan or textPath node, and the words inside all of them came to 0.095 percent of the 1,355,047 words of ordinary text on the same pages.
- 773 of the 1,074 home pages carried at least one inline SVG, a median of 23 on the pages that had any and 766 on spearmintlove.com, and 35,263 of the 37,019 elements carried no text of any kind.
- evershop.io is the one page where the choice costs something real: 488 words sat inside SVG text nodes against 911 words outside them on 6 October 2026, including product names, prices and article copy, and a second request the same day returned the same counts.
- Lantad's own extractor drops svg subtrees before it counts a word, in core/src/extract.ts, while this site's own flow diagrams put their labels in SVG text nodes, and those labels reach an extractor only because a visually hidden paragraph repeats them outside the graphic.
- Google's documentation on indexable file types, last updated 3 February 2026, lists Scalable Vector Graphics among the text-based formats it can index, and the accessible naming layer on these pages is separate again: 4,796 words sat in SVG title and desc elements and 2,217 more in aria-label attributes, which is not text content at all.
| What was counted | Figure | Against |
|---|---|---|
| Home pages read | 1,074 | of 1,419 hostnames asked |
| Pages carrying at least one inline SVG | 773 | 72.0 percent of the 1,074 |
| Inline SVG elements in total | 37,019 | median of 23 per page that had any |
| Elements carrying visible text nodes | 80 | 0.22 percent of the 37,019 |
| Elements carrying a title or desc | 1,676 | 4.5 percent of the 37,019 |
| Elements carrying no text of any kind | 35,263 | 95.3 percent of the 37,019 |
| Words inside visible SVG text nodes | 1,285 | of 1,355,047 page words, 0.095 percent |
| Words inside title and desc elements | 4,796 | not text content, read by name only |
Do AI crawlers read SVG text?
The question splits the way these questions usually do, and only one half is answerable from outside the crawler. Whether GPTBot or ClaudeBot keeps the characters inside an SVG when it reduces a page to text is a property of software this scan never touched, and no vendor documents it. Nothing here claims otherwise.
What is documented is that the characters are genuinely text rather than decoration. The W3C's SVG 2 specification, in its chapter on the text element, defines it as "a graphics element consisting of text" and says the character data inside text and tspan elements "define the glyphs to be rendered". That document was a Candidate Recommendation Snapshot carrying the date 6 October 2026 when it was read for this post. MDN's reference for the same element describes it as drawing a graphics element consisting of text, selectable and searchable in the way ordinary text is. Google is more direct still: its documentation on file types it can index, last updated 3 February 2026, lists Scalable Vector Graphics among the text-based formats, alongside HTML, XML and plain text, under the opening line that Google "can index the content of most text-based files".
So a word inside an SVG is a word, and at least one major search crawler treats a whole SVG file as an indexable text document. That makes the drop rule a real choice with a real cost, and the only way to size the cost is to count it. The answerable half of the question is therefore how much text is in there at all, which is the same move the run on image alt attributes made: if the supply is negligible, what any individual extractor does with it stops mattering. On this corpus the supply is negligible almost everywhere, and on a handful of pages it is most of the page.
What 37,019 inline graphics on 1,074 home pages carried
Inline SVG is close to universal and almost entirely mute. 773 of the 1,074 pages carried at least one, which is 72.0 percent, and 301 carried none. The pages that had any carried a median of 23 and a mean of 47.9, with a long tail: spearmintlove.com served 766 inline SVG elements, visitbritain.org 694, u.ae 393 and daytona.io 392. Not one of those four put a single visible word inside any of them.
That is the pattern across the corpus. Of the 37,019 elements, 80 carried a visible text node and 1,676 carried a title or desc child, leaving 35,263, or 95.3 percent, carrying no characters whatsoever. The overwhelming use of inline SVG on a 2026 home page is the icon: a chevron, a logo mark, a social glyph, a checkmark in a feature list, expressed as path data with no text in it. Path data is coordinates, so there is nothing for an extractor to lose.
The funnel behind the 1,074 is worth stating in full, because a rate computed on answering hosts is a statement about answering hosts. 1,419 hostnames were asked. 1,064 returned a robots.txt that parsed as plain text with HTTP 200, and 13 of those disallow LantadBot at the site root, so they were left alone: amsterdam.nl, wa.gov, seoul.go.kr, helsinki.fi, sciencedirect.com, pennmedicine.org, eluniversal.com.mx, theregister.com, sap.com, instacart.com, gmarket.co.kr, coralvilleanimalhospital.com and botcity.dev. The terms this crawler works under are on the scanner's bot page. Of the 1,406 home pages then requested, 33 returned no status at all, 23 of those failing at the network or TLS layer and 10 timing out at 20 seconds, and 299 answered with something other than HTTP 200 and HTML: 216 answered 403, 55 answered 503, 10 answered 202, 8 answered 429 and 5 answered 404. The 403s are the usual shape of a site refusing an unrecognised user agent, which this blog has measured head on when 92 of 1,056 sites refused GPTBot their robots.txt.
One deviation from the production scanner is worth naming rather than burying. The shipped pipeline treats a robots.txt that fails to resolve as disallow-all until it recovers; this run only skipped hosts whose robots.txt parsed and carried an explicit root disallow, so 355 hostnames with no parseable robots.txt were requested anyway. That makes this run slightly more permissive than a real scan, it affects which pages are in the denominator rather than any per-page figure, and the standing method for a single scan is on the methodology page. You can watch the same request against one host with what GPTBot sees.
Flow: 1,419 hostnames to Ask robots.txt as LantadBot; Ask robots.txt as LantadBot to 13 disallow the root; Ask robots.txt as LantadBot to 1,406 home pages requested; 1,406 home pages requested to 33 no status; 1,406 home pages requested (216 answered 403) to 299 not 200 HTML; 1,406 home pages requested to 1,074 read as raw bytes; 1,074 read as raw bytes (37,019 elements) to 773 carried inline SVG; 773 carried inline SVG (1,285 words) to 31 carried SVG text.
The pages where the drop rule costs real words
An average of 0.095 percent is not an argument for ignoring the problem, because the words are not spread evenly. 31 of the 1,074 pages carried any visible SVG text at all, 17 carried ten words or more, 6 carried fifty or more, and 3 carried a hundred or more. On four pages the text inside SVG amounted to more than a tenth of the ordinary text on the same page. All six of the largest cases were requested a second time the same day and returned identical counts.
evershop.io is the case that matters, and it is instructive because the words are not labels. 488 words sat inside SVG text nodes against 911 words of ordinary page text. Reading them back shows what they are: a vector mockup of the product's own admin interface and storefront, with the navigation of that mockup ("Dashboard", "Products", "Categories", "Collections", "Orders", "Customers", "Coupons"), product names and prices ("Dunk Low Retro Now PL", "$29.09", "$189.58", "minus 12 percent"), and whole sentences of article copy, including "We explore, discover, and express our unique style through the acquisition of extraordinary possessions" and "From sleek sneakers to elegant heels, 2024 is shaping up to be a year of bold and innovative shoe designs". A tool reading this page as text sees roughly a third less of it than a person does, and what it loses is the part that demonstrates what the product is.
The other five are smaller, and they split into two shapes. Four are technical pages where the words label a diagram or a rebuilt screenshot: uipath.com held 166 words in a single SVG against 1,419 outside it, sourcegraph.com 143 across 96 elements, cloudsmith.com 62 and atlasgo.io 59 across 151 elements. That is why the two strata accounting for most of the corpus total are static documentation sites, at 553 words, and SaaS, at 419. The fifth is not technical at all: etobicokerehab.com, a clinic on a visual site builder, set 93 words of its own service list as vector text, so the strings CHIROPRACTIC ADJUSTMENTS, ACUPUNCTURE, PHYSIOTHERAPY, MASSAGE THERAPY and THERAPEUTIC YOGA are in the response and in none of its ordinary text. The news stratum is the clean counter-example and the most telling one: 60 news home pages carried 3,520 inline SVG elements and exactly zero visible SVG words. Sites whose entire business is prose keep their prose in prose.
That distribution is the practical reading. A marketing site built from a visual builder mostly will not hit this, and where it does the words are short promotional set pieces rather than explanation: the six Wix and Squarespace pages carrying any between them reached 141 words, 93 of them the clinic service list above and the rest strings such as Voted 2025 Best of Inland Empire on skinbygabby.com and SUBSCRIBE & SAVE on jamcanjuice.com. A developer-facing page that ships its architecture diagram as inline SVG, which is a sensible choice for a crisp diagram at any zoom, can quietly put a meaningful share of its explanation somewhere an extractor may not look. The same asymmetry applies as with iframes, where 17 frames of 834 held fifty words or more: the aggregate is tiny and the individual cases are not. If your page is one of them, a stack guide is more use than a corpus average, and the risk compounds where JavaScript already supplies part of the prose.
| Host | Words in SVG | Ordinary page words | SVG as share of page text |
|---|---|---|---|
| evershop.io | 488 | 911 | 53.6 percent |
| cybozu.co.jp | 7 | 51 | 13.7 percent |
| sourcegraph.com | 143 | 1,087 | 13.2 percent |
| uipath.com | 166 | 1,419 | 11.7 percent |
| cloudsmith.com | 62 | 1,196 | 5.2 percent |
Our own scanner drops every one of them
The rule is not buried. core/src/extract.ts opens with its own rules in a comment, and the third line reads "drop entirely: script, style, noscript, template, svg, iframe subtrees". The set is one constant, DROP_TAGS, holding those six tag names, and an svg subtree is discarded before a single word is counted. Every figure this site has ever published about word counts, including the 1,355,047 words in this post, was produced by a counter that behaves this way.
On the evidence above the choice is defensible and this post is not arguing against it. 0.095 percent of the corpus text, concentrated in a handful of developer pages, bought in exchange for never mistaking icon path data or a chart axis for prose, is a trade most extractors would make. Dropping the whole subtree is also what makes the rule cheap and predictable, which matters for a scoring function that has to be stable across re-scans. Saying so is the point: it is a decision, and anyone reading an AI visibility score from this product or any other is reading that vendor's extraction decisions as much as their own markup.
What is awkward is where the rule points next. The diagrams in this post, and the one above, are produced by this site's own component library, and site/src/lib/visual.ts explains the flow diagram's design in a comment: it is "laid out server-side into inline SVG (no browser, no client JS), so a JS-blind AI crawler reads the text in the diagram and the plain-text summary beside it". The first half of that sentence describes exactly the markup Lantad's own extractor throws away. Those node labels are text nodes inside an svg element, so the scanner this company sells would score them at zero.
The design survives on its second half, and the detail is worth having because it is the fix. In site/src/components/ContentVisual.astro the flow diagram is followed by a paragraph element carrying the same sequence as ordinary prose, generated by flowSummary, and the stylesheet comment above its class says what it is for in as many words: "Visually hidden but present for assistive tech and text extraction". That paragraph is real text content outside the graphic, so it survives the drop, and because this scan fetched no stylesheets it could not have known the paragraph was visually hidden even if it had wanted to. The belt is the SVG text and the braces are the hidden paragraph, and on our own rules only the braces hold. That is the same class of finding as the shadow DOM text the extractor cannot reach and the fourteen noscript elements that held no text: what reaches an extractor is decided by markup shape, not by whether a person can read it. It also sits next to the run on hidden text, where 121,312 of 1,269,054 words were marked hidden from somebody, from the other direction: there the words counted and arguably should not have, here they do not count and arguably should.
-
text, tspan, textPathDropped Characters a person reads on the page. 80 of 37,019 elements carried any, 1,285 words in total. -
title, desc childrenDropped The accessible name and description. 1,676 elements carried one, 4,796 words in total. -
aria-label attributeNever counted An attribute value rather than text content, so no extractor counting text nodes sees it. 950 elements, 2,217 words. -
path, circle, rect dataCorrectly dropped Coordinates, not language. 35,263 of the 37,019 elements carry only this and lose nothing. -
A paragraph outside the svgCounted The route this site's own flow diagrams rely on, via a visually hidden paragraph repeating the sequence.
What to check, and what this run did not measure
The instruction that follows is small, which is the right size for a finding this concentrated. If a page explains itself through a diagram, and that diagram is inline SVG with its labels in text nodes, then the words in it are at the mercy of each extractor's drop list and at least one well-known extractor discards them. Repeating the content outside the graphic costs a paragraph. A figure element with a real figcaption, or a short prose summary next to the diagram, puts the same information where every extractor already looks, and it is the same answer the accessibility guidance has given for years for an unrelated reason. Checking it takes one request: fetch your own page with curl, strip the svg subtrees, and see whether the remaining text still explains the product.
The accessible naming layer deserves its own line, because it is larger than the visible one and is not text at any point. 1,676 of the 37,019 elements carried a title or desc child, holding 4,796 words, which is more than three times the visible SVG text in the corpus. Another 950 elements carried an aria-label, holding 2,217 words, and 3,018 elements declared role="img" without necessarily naming themselves at all. An aria-label is an attribute value, so a parser that walks text nodes never encounters it however carefully it is written, which is the same trap as links named only by an aria-label and the reason alt text has to be read as an attribute deliberately. Writing a good accessible name is worth doing. It is not a substitute for prose, and on these pages 159 of 1,074 sites have put a few words of description somewhere no text extractor will find them.
Five things this run did not measure, stated because the title asks more than the method can answer. No AI crawler fetched anything: every request came from LantadBot, so nothing here is a claim about what GPTBot, ClaudeBot or PerplexityBot retain from an SVG, and none of them document it. No JavaScript was executed and no stylesheet was fetched, so an SVG injected after load was not seen and every count is a floor. Only the home page was read on each host, and a home page is not where architecture diagrams usually live, so a documentation or product page would plausibly score higher: the evershop.io and atlasgo.io results both point that way. Referenced SVG files, loaded through an img or object element rather than written inline, were not fetched at all, and those are the files Google's indexable-types list is really about, so the whole external-file case is outside this measurement. And the corpus is an editorial sampling frame rather than a random draw, the standing caveat on the crawlability study, so these rates describe 1,419 hostnames and nothing wider.
The summary is two findings pointing opposite ways, and both are worth keeping. Across a thousand real home pages, dropping SVG costs almost nothing, so a scanner that does it is not lying to you. On the pages where it costs something it costs a third of the page, those pages are disproportionately the technical ones whose authors most want to be understood by an answer engine, and the cheapest remedy is a paragraph. If you want the vocabulary behind any of this, structured data is the next thing to read, and the platform specific guidance on getting cited by ChatGPT is where this site sends people after a scan.
- Diagram labels repeated outside the graphic A figcaption or a short prose summary next to the diagram survives every drop list. This site's own flow diagrams depend on exactly this.
- The page still reads with svg removed Strip the svg subtrees from your own HTML and check the remainder explains the product. On evershop.io the remainder loses the product names and prices.
- No sentences living only in text nodes 80 of 37,019 elements in this corpus carried visible text, and the worst cases were whole paragraphs rebuilt as vector art.
- aria-label treated as accessibility, not content 950 elements carried one, holding 2,217 words. An attribute value is not text content and no text extractor counts it.
- Referenced SVG files checked separately Google lists Scalable Vector Graphics among indexable file types. An SVG loaded through an img element is its own document and was not measured here.
Lantad
Published .
A dropped subtree is the quietest way to lose words. Most of what this blog counts is content that never arrived: prose held back behind a render, a page that answered 403, a file that resolved to nothing. Inline SVG is the opposite case. The bytes are in the response, already delivered, already parsed by the browser, and an extractor still has to decide whether the characters inside a graphic are text. Lantad decides they are not, and that decision is one line of code.
Common questions
Do AI crawlers read text inside SVG?
No vendor documents it, so this post makes no claim either way about GPTBot, ClaudeBot or PerplexityBot. What Lantad measured on 6 October 2026 is the supply side and its own behaviour: across 1,074 home pages, 80 of 37,019 inline SVG elements carried a visible text node, holding 1,285 words against 1,355,047 words of ordinary page text, and Lantad's own extractor discards all of them.
Is it a mistake to put text in an SVG?
Not on its own. SVG text is real text by specification and Google lists Scalable Vector Graphics among the file types it can index. The risk is relying on it as the only copy of something: an extractor that drops svg subtrees, as Lantad's does, sees none of it, so repeat anything load-bearing in a caption or a paragraph outside the graphic.
How much text does dropping SVG actually lose?
On this corpus, 0.095 percent of the words. The loss is concentrated rather than spread: 31 of 1,074 pages carried any SVG text, and on evershop.io 488 words sat inside SVG against 911 outside it, including product names, prices and article copy.
Does an aria-label on an SVG help an AI crawler?
Not if the crawler counts text nodes, because an aria-label is an attribute value rather than text content. 950 elements in this corpus carried one, holding 2,217 words, and a further 1,676 elements carried a title or desc child holding 4,796 words. Both are worth writing for accessibility, and neither is a substitute for prose on the page.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.