BlogFindings
SearchAction schema: 231 of 619 home pages declared a search endpoint, and robots.txt closed 57 of them
Lantad asked all 1,419 hostnames in this repository's committed corpus for robots.txt on 28 September 2026, then read the home page of every host that allowed it with no JavaScript executed. 1,076 answered HTTP 200 with HTML and 619 of those carried at least one JSON-LD block that parsed. 231 of the 619 declared a SearchAction, 224 of the 231 hosts gave an address a machine could actually build, and 57 of those addresses point at a path the same site's robots.txt disallows.
This run counted how many home pages still ship one and then did the thing the field invites, which is use it. Lantad requested /robots.txt from all 1,419 hostnames in the two committed corpus seed files, evaluated the site root for LantadBot before asking for anything else, and read the home page of every host that allowed it as raw bytes with no JavaScript executed. 1,076 home pages answered HTTP 200 with HTML, 619 of those carried at least one JSON-LD block that parsed, and 231 of the 619 declared a SearchAction. For each declared address the term lantadprobe was substituted into the template slot, the resulting path was evaluated against that same site's robots.txt, and where the file allowed it the address was requested once. The markup was nearly always well formed. What sat at the other end of it very often was not.
In short
- SearchAction schema is still widely deployed while the only search feature it was written for no longer exists: of the 619 home pages carrying parseable JSON-LD on 28 September 2026, 414 declared a WebSite node and 231 hung a SearchAction from it, and schema.org reports the type in its 10M or more domains usage band from an August 2026 aggregation of Google web index data.
- Google announced the removal of the sitelinks search box on 21 October 2024 and began removing the visual element on 21 November 2024, stating in the same post that while site owners can remove the structured data, there is no need to do so.
- 57 of the 224 usable search addresses pointed at a path the site's own robots.txt disallows to this scanner, among them cnn.com, forbes.com, walmart.com, techcrunch.com, aljazeera.com and nsw.gov.au. All 57 were decided by the wildcard group and not one by a rule naming any crawler.
- Of the 167 addresses robots.txt allowed and this scanner then requested on 28 September 2026, 129 returned a page containing the probe term and 38 did not. 33 of the 38 answered HTTP 200 with no trace of it, 9 of the 167 answered a non-200 at all, being six 404s, one 403, one 410 from alibaba.com and one 429, and 4 of those nine named the term on the error page.
- 237 of the 239 SearchAction nodes carried a query-input property, which schema.org does not define on SearchAction, and 6 carried query, which it does. 68 of the 231 hosts delivered Yoast's yoast-schema-graph class name in the page, so on those the markup is template output rather than a decision anybody took.
| Stage | Hostnames | What it excludes |
|---|---|---|
| Corpus seed files | 1,419 | Nothing. Two committed sampling frames. |
| Returned a robots.txt status | 1,388 | 31 where the request never produced a status |
| Allowed LantadBot at the root | 1,356 | 12 by an explicit rule, 51 by a 5xx treated as disallow |
| Home page answered 200 with HTML | 1,076 | 244 that answered otherwise, 36 requests that threw |
| Carried parseable JSON-LD | 619 | 451 with no block, 6 whose only blocks failed to parse |
| Declared a WebSite node | 414 | 205 that carried schema and named no website |
| Declared a SearchAction | 231 | 183 that named a website and offered no search |
| Gave a usable address | 224 | 7 hosts whose only template could not be built |
What is SearchAction schema, and what was it built for?
The vocabulary and the convention are two different things, and almost everything deployed on the web is the convention. schema.org/SearchAction defines the type in six words, "The act of searching for an object", puts it under Action, and defines exactly one property of its own on it: query, a sub property of instrument holding the query used on the action. The address itself belongs to a second type. schema.org/EntryPoint carries urlTemplate, described as an url template following RFC 6570 that will be used to construct the target of the execution of the action. Both pages, read on 28 September 2026, report the type in the 10M or more domains usage band, attributed to monthly aggregations of Google web index data for August 2026, which is the same dataset behind the finding that only sixteen schema.org types reach ten million domains.
What sites actually write is not that. 237 of the 239 nodes counted here carried a property called query-input, holding a string of the shape "required name=search_term_string" or an object of type PropertyValueSpecification with a valueName. schema.org does not define query-input on SearchAction at all. It comes from Google's own instructions for one search feature, and it is so dominant that the property the vocabulary does define is nearly absent: 6 of the 239 nodes carried query. The placeholder name follows the same pattern of copying rather than deciding. 215 of the templates used the slot name {search_term_string} that Google's example used, 11 used {search_term}, 8 used {query}, and one each used {srch_str}, {query_string} and {q}.
The feature is gone. On 21 October 2024, in a post titled Farewell, Sitelinks Search Box, Google announced it was removing the visual element starting on 21 November 2024, citing ten years of declining usage and a wish to simplify the results page. The post states that the change applies globally across all search results, in all languages and countries, that it "doesn't affect rankings or the other sitelinks visual element", and that once the element stopped showing, the Search Console rich results report for it would be removed and the Rich Results Test would stop highlighting the markup. It is explicit about what site owners should do, which is nothing: "While you can remove sitelinks search box structured data from your site, there's no need to do so. Unsupported structured data like this won't cause issues in Search, and won't trigger errors in Search Console reports." The old documentation page for the markup no longer describes it either. Requested on 28 September 2026, the path that held it redirects to the Search documentation updates log and the entry there is the announcement. None of that makes the markup an error, which matters for how the rest of this post reads: what follows is not a list of mistakes, it is a census of what a widely copied pattern now points at.
| Part | Defined by | Seen on the 239 nodes |
|---|---|---|
| SearchAction type | schema.org, under Action | 239 nodes on 231 hosts |
| query | schema.org, the only property on the type | 6 of 239 |
| query-input | Google's feature documentation, not schema.org | 237 of 239 |
| target as an EntryPoint object | schema.org | 115 of 239 |
| target as a bare string | schema.org, url expected | 123 of 239 |
| urlTemplate, an RFC 6570 template | schema.org, on EntryPoint | 234 absolute, 4 relative, 1 absent |
How many home pages still declare SearchAction schema?
231 of the 619 home pages carrying parseable JSON-LD, which makes it one of the most common things this corpus has been measured for. The useful denominator is narrower. A SearchAction normally hangs off the node describing the site as a whole, and 414 of the 619 pages declared a WebSite node, so 231 of 414 sites that described themselves at all also published a search address. 230 of the 239 nodes hung from a node typed WebSite, one of those on healthdirect.gov.au from a parent typed both Organization and WebSite. 3 hung from an Organization alone, on forbes.com, chewy.com and trendyol.com, one from a ProfessionalService on misa.vn and one from an OnlineStore on shein.com. The last 4 hung from a parent whose type was spelled Website, on chronicle.com, webnames.ca, ssense.com and gldn.com, which is not a type schema.org defines. Seven hosts declared more than one SearchAction, trendyol.com three.
For scale against earlier runs on these same hostnames, 412 of 1,080 home pages named their site in neither source Google reads first and 313 of 615 put an @id on no node at all. Set beside those, a search endpoint on 231 pages is not a neglected corner of the vocabulary. It is more common in this corpus than a declared language, since 186 of 615 pages carried an inLanguage value.
The distribution says plainly where it comes from. 33 of the 38 readable WordPress strata sites declared one, the highest rate of any stratum by a distance, against 2 of 44 single page application startups and 1 of 30 Framer sites. Rather than infer the cause, it was measured: 68 of the 231 hosts delivered the class name yoast-schema-graph in the page, which is the class Yoast SEO puts on the script tag holding its schema graph, and on cambridgeday.com that is verifiably the tag carrying the SearchAction. 97 of the 231 delivered a wp-content or wp-json path, and 17 carried a Rank Math marker. So on at least 68 of these 231 sites nobody chose to publish a search endpoint. A plugin did, which is the same mechanism behind WordPress sites whose robots.txt names no AI crawler: the defaults are the policy. That also explains why the pattern survived the feature. Deleting markup nobody added is not a task anyone is assigned.
Does the address in the markup actually answer a search?
This is the question the field exists to answer and the reason this run is not another adoption count. A urlTemplate is a promise a machine can test, so it was tested. 224 of the 231 hosts gave an absolute address with exactly one slot to fill, which is the minimum needed to build a request without guessing. The 7 that did not are worth naming individually because each fails differently: utoronto.ca declared a SearchAction with no target at all; humanitas.it, gap.com and coalatree.com gave a relative template, and coalatree.com's is the placeholder on its own with no path around it; nike.com put the same slot in twice, once as the search term and once in an advertising parameter; france.fr wrote ?query=search_term_string with no braces, so there is no slot; and kempinski.com percent-encoded its braces, which makes the literal characters %7B and %7D part of the query string rather than a template. Neither of the last two is an RFC 6570 template.
For the 224, lantadprobe was substituted for the slot and the resulting path was evaluated against the site's own robots.txt with the scanner's shipped matcher before any request was made, exactly as the scanner's conduct policy requires. 57 were disallowed and were never requested, which is the subject of the next section. 9 pointed at a different hostname, and for those the target host's own robots.txt was not fetched, so those nine were requested on the strength of the declaring site's file alone and that is a limit of this run rather than a result. The remaining 167 were requested once.
129 of the 167 returned a body containing the string lantadprobe, which is the test that the address is a search endpoint rather than merely a URL: a search results page normally repeats the term, including when it finds nothing. 125 of those 129 carried HTTP 200 and 4 were error pages that named the term anyway. That leaves 38 of the 167 that produced no evidence of a search: 33 answered HTTP 200 with no trace of the term, and 5 answered a non-200 that did not name it either. 9 of the 167 answered a non-200 in all. Those 33, together with all 9 non-200 responses, 42 requests in total, were asked a second time later the same day alongside a fresh copy of the home page. 39 of the 42 reproduced their result exactly. handelsblatt.com and cambridgeday.com did not, and both had been throttled on the first pass, cambridgeday.com with an outright 429: on the recheck each returned a full page echoing the term. Counting the better of the two attempts, 131 of the 167 answered a search and 36 did not.
Flow: 231 hosts declared a SearchAction to Absolute template with one slot?; Absolute template with one slot? (no) to 7 hosts unusable; Absolute template with one slot? (yes, 224) to Path allowed by the site's own robots.txt?; Path allowed by the site's own robots.txt? (disallowed) to 57 never requested; Path allowed by the site's own robots.txt? (allowed) to 167 requested once; 167 requested once to 129 echoed the probe term; 167 requested once to 33 answered 200, no term; 167 requested once to 5 answered a non-200, no term.
57 sites published a search endpoint their own robots.txt closes
This is the finding that is hard to read as anything but an accident, because two files on the same host say opposite things about the same address. The markup says: search me here. The robots.txt says: not there. 57 of the 224 usable addresses, a quarter of them, sit on a path the same site disallows to this scanner, and the sites doing it are not obscure. cnn.com, nbcnews.com, cbsnews.com, forbes.com, aljazeera.com, irishtimes.com, techcrunch.com and the National Geographic site all publish a search endpoint in their markup and close it in their rules. So do walmart.com, chewy.com, johnlewis.com, marksandspencer.com, selfridges.com, trendyol.com and coolblue.nl. So do nsw.gov.au, texas.gov, mcgill.ca, plannedparenthood.org and stanfordhealthcare.org.
Not one of the 57 was decided by a group naming a crawler. All 57 were decided by the group headed with an asterisk, which is the same pattern as the finding that 559 of the 581 paths GPTBot lost were closed by a wildcard rule that never named it. RFC 9309 is why that is the expected outcome rather than a surprise: a crawler takes the most specific group that matches its own product token and falls back to the wildcard, so a site that never writes a rule for any AI crawler hands all of them whatever the wildcard says, including the rule that closes its search page.
The rules themselves divide into two intentions. 37 of the 57 were decided by a pattern that names a search path, which is a deliberate decision about the search page: Disallow /search on cnn.com and walmart.com, Disallow /search/ on aljazeera.com and forbes.com, Disallow */search/* on cbsnews.com, and the localised equivalents, Disallow /recherche on lapresse.ca, Disallow /buscar? on eltiempo.com, Disallow /ricerca/ on ansa.it, Disallow /nl/zoeken?q=* on wur.nl and Disallow */suche on visitberlin.de. The other 20 were decided by a broader parameter block that was almost certainly aimed at faceted crawling and caught the search page on the way past: Disallow /*?* on nsw.gov.au, busy.in, mekari.com and reliancedigital.in, Disallow /?* on patsnap.com, Disallow *query=* on chewy.com, Disallow *?freeText=* on selfridges.com. 12 of the 57 are the WordPress query shape, where the declared endpoint is /?s= and the file blocks exactly that path. Anybody wanting to check their own two files against each other can put the path into the robots.txt tester; the more general point is that a robots.txt can be perfectly well formed and still contradict the page it sits beside, which no validator that reads it in isolation can tell you.
| Host | Endpoint declared in the markup | Rule that closed it |
|---|---|---|
| cnn.com | /search?q= | disallow /search |
| nbcnews.com | /search?q= | disallow /search |
| forbes.com | /search/?q= | disallow /search/ |
| aljazeera.com | /search/ | disallow /search/ |
| cbsnews.com | /search/?q= | disallow */search/* |
| walmart.com | /search?q= | disallow /search |
| johnlewis.com | /search?search-term= | disallow /search* |
| chewy.com | /s?query= | disallow *query=* |
| nsw.gov.au | /search?q= | disallow /*?* |
| mcgill.ca | /search/?query= | disallow */search/ |
| techcrunch.com | /?s= | disallow /?s= |
| stanfordhealthcare.org | /search-results.html/ | disallow /search-results.html/ |
What the 33 silent 200s and the nine failures were
A 200 with no sign of the term is the more interesting half, because it looks like success to anything that checks status codes. Five of the 33 returned a body byte for byte identical to the home page fetched seconds later: thenationalnews.com, ovhcloud.com, mapfre.com, fibilaw.com and pollyreach.ai. On those the declared search address is, as served, the home page with a query string on it, which is the same class of answer as a site that returns HTTP 200 for a URL that does not exist. eugenemontessorischool.com answered 200 with a zero byte body, twice.
Most of the rest are almost certainly client side search. frontiersin.org returned 2,517 bytes against a 248,726 byte home page, abc.net.au 14,074 against 798,277, emiratesnbd.com 34,496 against 1,367,867: a shell that will fetch results once a browser runs its JavaScript, and therefore nothing at all to a crawler that does not. This is the same boundary measured directly when 11 of 271 pages returned no readable words until the bundle ran, and it cuts the same way here. No JavaScript was executed in this run, so a search results page assembled in the browser counts as no answer, which is a true statement about what a crawler is served and not a claim about what a person sees.
The nine non-200s were clearer. Six returned 404 for their own declared search address: usps.com, cloudflare.com, statefarm.com, tec.mx, scroll.in and evolvehealing.net, and the last three of those named the term in the 404 page they returned, so the search ran and the status did not say so. scmp.com returned 403 with the term in the body. alibaba.com returned 410 Gone, which is the one response in the set that says the address was deliberately retired while the markup advertising it was left in place. Three of the 224 templates also point at a hostname that is not just a subdomain but a different brand: spectator.co.uk points at spectator.com, highlandscurrent.com at highlandscurrent.org, and nowhabersham.com at nowgeorgia.com. Those may be deliberate, and the JSON-LD gives a reader no way to tell.
Two limits govern every figure above. Only JSON-LD was read, so a SearchAction expressed in microdata or RDFa is invisible here, and when 385 readable home pages in the platform corpus were parsed for all three syntaxes 22 carried microdata and one carried RDFa, so the undercount is real but small. And one page was read per hostname, the home page, so this describes home pages rather than sites.
Answered a search
- 125 returned HTTP 200 containing the probe term
- 4 named the term on an error page
- 2 more echoed it on a recheck after being throttled
- 157 of the 158 that answered 200 sent an HTML content type
Answered something else
- 33 returned HTTP 200 with no trace of the term
- 5 of those returned the home page byte for byte
- 1 returned a zero byte body, twice
- 5 more answered a non-200 without the term
Is SearchAction markup worth keeping for AI search?
Nothing measured here says it earns anything, and this post does not claim it does. No crawler vendor documents reading a SearchAction. The honest position is the one this site has taken about other retired features: markup outliving its consumer is not a defect, it is just markup nobody is reading, which is exactly what happened when the FAQ rich result went away and the FAQPage markup stayed. Google's own instruction is to leave it, and treating an unsupported property as an error is how AI visibility advice turns into busywork.
There is a narrower argument for keeping it that this measurement can support, and it is about agents rather than search features. An agent that wants to find something on a site has almost nothing to work with. When this corpus was asked for the one agent discovery path any standards registry lists, one site in 1,419 returned a usable agent card, and the browser-side alternative needs a browser to be discovered at all. A urlTemplate in the home page markup needs neither. It is a machine-readable address for site search, present on 231 of these hostnames, which is more than any purpose-built agent standard has managed. Against that, an AI crawler reaching for it meets the 57 robots.txt refusals and the 33 silent 200s counted above, so it is a promise this run could see kept on 131 of the 224 hosts whose address could be built.
The inconvenient part is ours. lantad.co declares a WebSite node with no potentialAction on it, because the site has no search to point at, so a reader arriving here from this post finds the thing it counts absent on the page counting it. That is the honest state and not an oversight worth hiding: the scanner's own scoring does not read potentialAction either, and what GPTBot sees will show you your own markup without any opinion about this field. If you want to know which tokens a rule would reach before you write one, the AI crawler reference lists what this scanner evaluates.
What this run establishes is narrow and worth stating exactly. On 28 September 2026, across 1,419 hostnames in a committed editorial sampling frame that is not a random sample of the web, 231 home pages published a machine-readable site search address, 224 of those addresses could be built, 57 were closed by the publisher's own robots.txt, and 131 answered a request with a page that repeated the term. Everything else is inference. Nobody has published what an answer engine does with this field, the same way nobody has published what one does with an aggregate rating or with the sitemap a robots.txt declares, and a measurement of markup is not a measurement of outcome.
-
Address usable224 of 231 hosts An absolute template with exactly one slot to fill. On 7 hosts no declared template could be built. -
Allowed by robots.txt167 of 224 57 point at a path the same site disallows, all 57 decided by the wildcard group. -
Answered the search131 of 167 Returned a page repeating the probe term on at least one of two requests. -
Answered something else36 of 167 After the recheck, 32 answered 200 with no trace of the term and 4 a non-200 without it, including one 410 Gone. -
Read by a named consumerNone documented The Google feature ended 21 November 2024 and no crawler vendor documents reading the field.
Lantad
Published .
A page can tell a machine what it is, who publishes it and when it changed. SearchAction schema is the one common pattern that tells a machine what it can do: here is an address, put a term in this slot, and you will get my search results. It is the closest thing most sites publish to an API, and it was written for a single consumer that stopped existing almost two years ago.
Common questions
Should I remove SearchAction schema from my site?
Google says there is no need. Its 21 October 2024 announcement removing the sitelinks search box states that while you can remove sitelinks search box structured data from your site, there is no need to do so, because unsupported structured data will not cause issues in Search and will not trigger errors in Search Console reports. The same post notes that site names use a variation of WebSite structured data that continues to be supported, so removing the surrounding WebSite node is a different and worse idea than removing the action inside it. The check worth doing instead is whether the address the markup declares still works, which on 57 of the 224 endpoints measured here it cannot, because the site's own robots.txt disallows it.
What is the difference between query and query-input on a SearchAction?
One is schema.org vocabulary and the other is a Google convention. schema.org defines query on SearchAction as a sub property of instrument holding the query used on the action, and defines no property called query-input anywhere on the type. query-input comes from Google's instructions for the sitelinks search box, where it named the required variable that the slot in urlTemplate refers to. In this sample of 239 nodes read on 28 September 2026, 237 carried query-input and 6 carried query, so almost every deployment on these hostnames follows the feature documentation rather than the vocabulary.
Does a sitelinks search box still appear in Google results?
No. Google began removing the visual element on 21 November 2024, applying the change globally across all search results in all languages and countries, and said the Search Console rich results report for it would be removed and the Rich Results Test would stop highlighting the markup. The documentation page that described the markup now redirects to Google's Search documentation updates log. The announcement also states the removal does not affect rankings or the other sitelinks visual element.
Can an AI crawler use the search endpoint declared in SearchAction schema?
Nothing published says any of them tries. No AI crawler vendor documents reading a SearchAction, so treating the field as an AI visibility signal would be a guess. Mechanically it is usable: 224 of the 231 hosts measured on 28 September 2026 gave an absolute template with one slot, and substituting a term produced a working search on 131 of the 167 that robots.txt allowed a request to. The obstacle is not the markup. It is that a quarter of the declared endpoints sit on a path the publisher disallows, so a crawler that respects robots.txt cannot follow the invitation the same site printed.
See what AI can read on your site
Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.