AI Crawlers vs Scanners: 40 of Our 1,797 404s Were a Search Engine
Disclosure
This article is promotional content produced as part of the Semrush Creator Accelerator programme. It contains an affiliate link to Semrush — if you buy through it we may earn a commission at no extra cost to you — and we may also receive a reward from Semrush for publishing it. Every measurement below is from our own infrastructure. Semrush did not supply, review or approve any of it.
The technical advice for AI visibility in 2026 is consistent and, as far as we can tell, correct. Serve real HTML. Do not hide content behind JavaScript. Add schema. Semrush's own trend brief puts numbers on the gap: roughly 30% of sites have critical content that only appears after JS executes, and only about 12% of domains implement any schema at all.
We are, by accident rather than virtue, the site that advice is trying to produce.
Every page on this site is prerendered at build time. Twenty-seven route files declare force-static; seven dynamic routes set dynamicParams = false so an unknown URL 404s at the edge rather than invoking anything. A tool page ships 1,956 words of content in the raw HTML, before a single line of JavaScript runs, carrying fifteen schema types — including the Organization and WebSite entities the brief specifically names.
So we pass the checklist. Comfortably. Which left us free to ask a different question, and the answer to that one was genuinely surprising.
We measured who was actually requesting our pages.
40 of 1,797
Over a 48-hour window, 1,797 requests to this domain returned a 404.
Forty of them came from a verified search crawler. That is 2.2%.
Roughly 547 were credential scanners — automated sweeps for .env files, backup archives, admin panels and API keys. The remaining 1,210 were Meta's crawlers and generic probes, including a steady stream hitting /_next/image, an endpoint we deliberately do not run.
We had been treating that 404 count as an SEO problem. It is not an SEO problem. It is a security-noise problem wearing an SEO problem's clothes, and the distinction matters because the two have completely different fixes.
The redirects we did not build
This nearly cost us real work. We had a plan to build 394 redirect rules for pruned pages, on the reasoning that something out there was still requesting them and 404s are bad.
Then we checked who. Almost the entire request volume for those URLs was Meta's crawler. We would have hand-maintained 394 rules, added them to a redirect table that already shadows pages when it grows careless, and served them overwhelmingly to a social-media crawler that was never going to send us a visitor.
We did not build them.
Before you build for a crawler, check which crawler is asking. It is one query against your edge logs and we have not seen it on a single technical-SEO checklist, including the ones we wrote.
Measure the wire, not the disk
The second finding is less dramatic and probably more generally useful, because it is a mistake that would have made us publish a false number.
We measured forty built tool pages two ways:
| uncompressed | brotli, as served | |
|---|---|---|
| Full page | 218.3 KB | 23.7 KB |
| With the hydration payload stripped | 92.7 KB | 14.2 KB |
| Share of the page that is hydration data | 58% | 40% |
The compression ratio is 9.2×, and it is not uniform. The React hydration payload is repetitive JSON, so it compresses far better than the prose and markup around it. On disk it looks like 58% of the page. On the wire it is 40%.
We were about to report that removing it would cut page weight by 56%. On the wire the real figure is about 33% — because the thing we proposed to remove was the thing that compressed best.
Quoting the uncompressed number would not have been a lie. It would have been a measurement of something nobody experiences. No crawler, no browser and no AI fetcher ever receives the uncompressed page.
If you are going to quote a page-weight figure, quote the one that crosses the network. The gap between the two is not a rounding error; here it was the difference between a 56% claim and a 33% reality.
Six things we were wrong about
We keep a list of hypotheses that are dead, with the evidence that killed them, specifically so nobody reopens them in three months. It is the most useful document in the investigation.
| What we believed | Verdict |
|---|---|
| Client-side navigation requests were the hidden cost | Dead — 22 requests and 0.4 MB across 72 hours. A rounding error. |
| Our internal links point at pruned pages | Dead — 344 comparison links, 74 category links and 485 tool-page links all resolve. Zero broken. |
| The 404s are hurting us with search engines | Dead — 40 of 1,797 came from a verified crawler. |
| Traffic growth explains the cost | Dead — flat at roughly 27,500 a day for nineteen days. |
| HTML was not being cached properly | Mostly wrong — the hit rate on tool pages is 87.8%. |
| Caching accounted for ~36% of the overage | Wrong, published internally, and retracted. |
That last row is the honest one. We asserted a specific percentage, acted on it briefly, and then found it was not supported. It is in the list because a killed hypothesis that quietly disappears gets re-proposed by the next person, including when the next person is you.
One more, which is the one worth copying: the single largest uncached item on the whole domain turned out to be the automatically generated social preview image — 174 requests pulling 31.1 MB across 72 hours, at roughly 178 KB each. Nothing in any audit we had run thought to look at it, because it is not a page.
What we still cannot explain
The investigation has an open hole and we are not going to paper over it.
Modelling the cost from request volume and page weight produces a figure that undershoots the observed number by about 1.5×. Something is consuming requests that we have not identified. Until that reconciles, every savings estimate in our own working document carries roughly a 50% error bar, and it says so at the top.
We mention it because this is the normal state of a technical audit and almost nobody publishes it. A clean narrative where every number reconciles is usually a narrative where somebody stopped checking early.
What these numbers do not show
- This is one domain's request log over 48 hours. Crawler mix varies enormously by niche, backlink profile and age. Your 2.2% will not be our 2.2%.
- 404s from scanners are not harmless, they are just not an SEO problem. They are a reason to rate-limit, not a reason to build redirects.
- We have not shown that passing the readability checklist is unnecessary. We pass it, so we cannot measure what failing it would cost us. The brief's 30% and 12% figures are theirs, not ours, and we have not reproduced them.
- Being readable is necessary and it is not sufficient. We are read constantly — roughly eighty AI fetches a day — and cited rarely. Readability gets you into the corpus. It does not get you into the answer.
Run this on your own site
- Pull your edge log and group 404s by user agent, separating verified crawlers from everything else. If you have Cloudflare, this is a dashboard view, not a project.
- Before building redirects for a pruned section, check who requests it. The answer is often a social crawler or a scanner.
- Quote page weight after compression. Measure both, and notice when the gap between them changes which component is worth removing.
- Audit the things that are not pages — generated preview images, feeds, manifests. They escape page-level audits by not being pages.
- Keep a killed-hypotheses list with its evidence. It is the cheapest document you will ever maintain and it stops the same investigation happening twice.
- Write your error bar next to your estimate. If the model does not reconcile, say by how much.
- Check your raw HTML contains your content. Strip the scripts and count the words. If the number is small, nothing above matters yet.
Closing the loop between a fix and its effect
Everything above was assembled from edge logs, a build output directory and a spreadsheet. That is fine for a one-off investigation, and it is exactly where the method runs out: we can tell you our crawl surface changed, and we cannot tell you whether any of it moved our presence in AI answers, because those two things live in different systems and we joined them by hand on one date.
That join is the gap Semrush One is positioned to close. Semrush's home page, checked 7 October 2026, puts its job as helping "your brand show up where customers search, ask, and decide", and their AI visibility feature page lists Prompt Tracking for "daily AI visibility updates" alongside an AI-Cited Media tool for the outlets LLMs cite most. Their Creator Accelerator brief adds the piece our own audit never ran: that Site Audit crawls with JavaScript rendering enabled and reports the difference between raw HTML and the rendered page.

That panel names the question our own audit never asked. We measured which crawlers were requesting pages that do not exist. We never measured whether the crawlers we want are being blocked from the pages that do — and a site quietly refusing ChatGPT-User would produce zero citations for a reason no amount of schema would fix.
We have not used Semrush One. We are a Semrush affiliate, not a customer, and the description above is of Semrush's documentation rather than a review. Their pricing page states "Try Semrush free for seven days. Cancel anytime" — checked 7 October 2026. The screenshot below is from 26 September, when that offer sat on Semrush's home page under the headline "Win in search, era after era". That headline now reads "Be found everywhere search happens" — a fair reminder that a dated capture is evidence for its date and nothing after it. We have not taken the trial, so read the terms as Semrush's claim rather than our test.

Semrush One
We can tell you our crawl surface changed. We cannot tell you whether it moved our presence in AI answers, because we joined those two systems by hand, once.
See Semrush OneWhat we actually learned
The readability checklist is not wrong. It is just finished sooner than people expect, and then it stops telling you anything.
We serve static HTML with fifteen schema types and two thousand words before any JavaScript runs. Having done that, the checklist had no further questions — while 97.8% of our 404 traffic came from things that will never send us a reader, the biggest uncached object on the domain was not a page, and the page-weight figure we were about to publish was measuring bytes nobody receives.
The checklist tells you how to be readable. The request log tells you who is reading. They are different instruments and only one of them is in the standard kit.
Sources and method
| Figure | Source | Date |
|---|---|---|
| 1,797 404s in 48 hours; 40 from verified search crawlers; ~547 credential scanners | Cloudflare edge logs for our own domain | 2026-09-20 window |
394 /alternatives/ redirects rejected after checking requester identity | Same logs; our own decision record | 2026-09-20 |
| 218.3 KB uncompressed vs 23.7 KB brotli across 40 built tool pages; 9.2× ratio | Our own build output, measured both ways | 2026-09-20 |
| Hydration payload 58% uncompressed / 40% on the wire; 56% vs 33% saving | Same measurement | 2026-09-20 |
| Social preview image: 174 requests, 31.1 MB per 72h | Cloudflare edge logs | 2026-09-20 |
| 87.8% cache hit rate on tool pages; ~27,500 requests/day flat for 19 days | Cloudflare analytics | 2026-09-20 |
| Unresolved 1.5× gap; ~50% error bar on savings estimates | Our own diagnosis document, stated at the top of it | 2026-09-20 |
| 27 force-static route files; 7 with dynamicParams false; 1,956 words in raw HTML; 15 schema types | This repository and its built output | 2026-10-07 |
| ~30% of sites with JS-dependent content; ~12% implementing schema | Semrush's trend brief — not verified by us | as published |
| Semrush One and AI visibility wording, verbatim | Semrush's own home page and AI visibility feature page | 2026-10-07 |
| Site Audit crawling with JavaScript rendering enabled | Semrush's Creator Accelerator brief | retrieved 2026-09-22 |
| "Try Semrush free for seven days. Cancel anytime" | Semrush's pricing page | 2026-10-07 |
| Screenshot wording "Win in search, era after era" (headline since changed) | Semrush's home page, captured by us | 2026-09-26 |
Written by Sohail Akhtar. Figures describe one domain on the dates given and should be read as one domain's measurements, not as norms.
Some links may be affiliate links. We may earn a small commission at no extra cost to you.
Related Articles
Explore more guides and reviews from our experts.