Skip to content

AI Crawlers vs Scanners: 40 of Our 1,797 404s Were a Search Engine

Sohail Akhtar
Researched by Sohail AkhtarTheToolsVerse
October 7, 202610 min read
Of 1,797 requests that returned 404 across 48 hours, only 40 came from verified search crawlers. Roughly 547 were credential scanners and the remaining 1,210 were Meta crawlers and automated probes.

Disclosure

This article is promotional content produced as part of the Semrush Creator Accelerator programme. It contains an affiliate link to Semrush — if you buy through it we may earn a commission at no extra cost to you — and we may also receive a reward from Semrush for publishing it. Every measurement below is from our own infrastructure. Semrush did not supply, review or approve any of it.

The technical advice for AI visibility in 2026 is consistent and, as far as we can tell, correct. Serve real HTML. Do not hide content behind JavaScript. Add schema. Semrush's own trend brief puts numbers on the gap: roughly 30% of sites have critical content that only appears after JS executes, and only about 12% of domains implement any schema at all.

We are, by accident rather than virtue, the site that advice is trying to produce.

Every page on this site is prerendered at build time. Twenty-seven route files declare force-static; seven dynamic routes set dynamicParams = false so an unknown URL 404s at the edge rather than invoking anything. A tool page ships 1,956 words of content in the raw HTML, before a single line of JavaScript runs, carrying fifteen schema types — including the Organization and WebSite entities the brief specifically names.

So we pass the checklist. Comfortably. Which left us free to ask a different question, and the answer to that one was genuinely surprising.

We measured who was actually requesting our pages.

Of 1,797 requests returning 404 across 48 hours, only 40 came from verified search crawlers. Roughly 547 were credential scanners probing for secrets, and the remaining 1,210 were Meta crawlers and automated probes.
Of 1,797 requests returning 404 across 48 hours, only 40 came from verified search crawlers. Roughly 547 were credential scanners probing for secrets, and the remaining 1,210 were Meta crawlers and automated probes.

40 of 1,797

Over a 48-hour window, 1,797 requests to this domain returned a 404.

Forty of them came from a verified search crawler. That is 2.2%.

Roughly 547 were credential scanners — automated sweeps for .env files, backup archives, admin panels and API keys. The remaining 1,210 were Meta's crawlers and generic probes, including a steady stream hitting /_next/image, an endpoint we deliberately do not run.

We had been treating that 404 count as an SEO problem. It is not an SEO problem. It is a security-noise problem wearing an SEO problem's clothes, and the distinction matters because the two have completely different fixes.

The redirects we did not build

This nearly cost us real work. We had a plan to build 394 redirect rules for pruned pages, on the reasoning that something out there was still requesting them and 404s are bad.

Then we checked who. Almost the entire request volume for those URLs was Meta's crawler. We would have hand-maintained 394 rules, added them to a redirect table that already shadows pages when it grows careless, and served them overwhelmingly to a social-media crawler that was never going to send us a visitor.

We did not build them.

Before you build for a crawler, check which crawler is asking. It is one query against your edge logs and we have not seen it on a single technical-SEO checklist, including the ones we wrote.

Measure the wire, not the disk

The second finding is less dramatic and probably more generally useful, because it is a mistake that would have made us publish a false number.

We measured forty built tool pages two ways:

uncompressedbrotli, as served
Full page218.3 KB23.7 KB
With the hydration payload stripped92.7 KB14.2 KB
Share of the page that is hydration data58%40%

The compression ratio is 9.2×, and it is not uniform. The React hydration payload is repetitive JSON, so it compresses far better than the prose and markup around it. On disk it looks like 58% of the page. On the wire it is 40%.

We were about to report that removing it would cut page weight by 56%. On the wire the real figure is about 33% — because the thing we proposed to remove was the thing that compressed best.

Quoting the uncompressed number would not have been a lie. It would have been a measurement of something nobody experiences. No crawler, no browser and no AI fetcher ever receives the uncompressed page.

If you are going to quote a page-weight figure, quote the one that crosses the network. The gap between the two is not a rounding error; here it was the difference between a 56% claim and a 33% reality.

Six things we were wrong about

We keep a list of hypotheses that are dead, with the evidence that killed them, specifically so nobody reopens them in three months. It is the most useful document in the investigation.

What we believedVerdict
Client-side navigation requests were the hidden costDead — 22 requests and 0.4 MB across 72 hours. A rounding error.
Our internal links point at pruned pagesDead — 344 comparison links, 74 category links and 485 tool-page links all resolve. Zero broken.
The 404s are hurting us with search enginesDead — 40 of 1,797 came from a verified crawler.
Traffic growth explains the costDead — flat at roughly 27,500 a day for nineteen days.
HTML was not being cached properlyMostly wrong — the hit rate on tool pages is 87.8%.
Caching accounted for ~36% of the overageWrong, published internally, and retracted.

That last row is the honest one. We asserted a specific percentage, acted on it briefly, and then found it was not supported. It is in the list because a killed hypothesis that quietly disappears gets re-proposed by the next person, including when the next person is you.

One more, which is the one worth copying: the single largest uncached item on the whole domain turned out to be the automatically generated social preview image — 174 requests pulling 31.1 MB across 72 hours, at roughly 178 KB each. Nothing in any audit we had run thought to look at it, because it is not a page.

What we still cannot explain

The investigation has an open hole and we are not going to paper over it.

Modelling the cost from request volume and page weight produces a figure that undershoots the observed number by about 1.5×. Something is consuming requests that we have not identified. Until that reconciles, every savings estimate in our own working document carries roughly a 50% error bar, and it says so at the top.

We mention it because this is the normal state of a technical audit and almost nobody publishes it. A clean narrative where every number reconciles is usually a narrative where somebody stopped checking early.

What these numbers do not show

  • This is one domain's request log over 48 hours. Crawler mix varies enormously by niche, backlink profile and age. Your 2.2% will not be our 2.2%.
  • 404s from scanners are not harmless, they are just not an SEO problem. They are a reason to rate-limit, not a reason to build redirects.
  • We have not shown that passing the readability checklist is unnecessary. We pass it, so we cannot measure what failing it would cost us. The brief's 30% and 12% figures are theirs, not ours, and we have not reproduced them.
  • Being readable is necessary and it is not sufficient. We are read constantly — roughly eighty AI fetches a day — and cited rarely. Readability gets you into the corpus. It does not get you into the answer.

Run this on your own site

  1. Pull your edge log and group 404s by user agent, separating verified crawlers from everything else. If you have Cloudflare, this is a dashboard view, not a project.
  2. Before building redirects for a pruned section, check who requests it. The answer is often a social crawler or a scanner.
  3. Quote page weight after compression. Measure both, and notice when the gap between them changes which component is worth removing.
  4. Audit the things that are not pages — generated preview images, feeds, manifests. They escape page-level audits by not being pages.
  5. Keep a killed-hypotheses list with its evidence. It is the cheapest document you will ever maintain and it stops the same investigation happening twice.
  6. Write your error bar next to your estimate. If the model does not reconcile, say by how much.
  7. Check your raw HTML contains your content. Strip the scripts and count the words. If the number is small, nothing above matters yet.

Closing the loop between a fix and its effect

Everything above was assembled from edge logs, a build output directory and a spreadsheet. That is fine for a one-off investigation, and it is exactly where the method runs out: we can tell you our crawl surface changed, and we cannot tell you whether any of it moved our presence in AI answers, because those two things live in different systems and we joined them by hand on one date.

That join is the gap Semrush One is positioned to close. Semrush's home page, checked 7 October 2026, puts its job as helping "your brand show up where customers search, ask, and decide", and their AI visibility feature page lists Prompt Tracking for "daily AI visibility updates" alongside an AI-Cited Media tool for the outlets LLMs cite most. Their Creator Accelerator brief adds the piece our own audit never ran: that Site Audit crawls with JavaScript rendering enabled and reports the difference between raw HTML and the rendered page.

Semrush One's own product page, captured 26 September 2026: a "Website optimization" panel showing Site health at 80% beside AI Search health at 44%, with crawled-page counts and a "Blocked from AI Search" list naming Claude-SearchBot, Perplexity-User and ChatGPT-User. This is Semrush's illustration of the product, not our own account.
Semrush One's own product page, captured 26 September 2026: a "Website optimization" panel showing Site health at 80% beside AI Search health at 44%, with crawled-page counts and a "Blocked from AI Search" list naming Claude-SearchBot, Perplexity-User and ChatGPT-User. This is Semrush's illustration of the product, not our own account.

That panel names the question our own audit never asked. We measured which crawlers were requesting pages that do not exist. We never measured whether the crawlers we want are being blocked from the pages that do — and a site quietly refusing ChatGPT-User would produce zero citations for a reason no amount of schema would fix.

We have not used Semrush One. We are a Semrush affiliate, not a customer, and the description above is of Semrush's documentation rather than a review. Their pricing page states "Try Semrush free for seven days. Cancel anytime" — checked 7 October 2026. The screenshot below is from 26 September, when that offer sat on Semrush's home page under the headline "Win in search, era after era". That headline now reads "Be found everywhere search happens" — a fair reminder that a dated capture is evidence for its date and nothing after it. We have not taken the trial, so read the terms as Semrush's claim rather than our test.

Semrush One's own product page: "Win in search, era after era," with a "Try free for 7 days" button and, beneath it, "Unlimited access to all Semrush One tools." Captured from Semrush's page on 26 September 2026 — we have not taken the trial.
Semrush One's own product page: "Win in search, era after era," with a "Try free for 7 days" button and, beneath it, "Unlimited access to all Semrush One tools." Captured from Semrush's page on 26 September 2026 — we have not taken the trial.

Semrush One

We can tell you our crawl surface changed. We cannot tell you whether it moved our presence in AI answers, because we joined those two systems by hand, once.

See Semrush One

What we actually learned

The readability checklist is not wrong. It is just finished sooner than people expect, and then it stops telling you anything.

We serve static HTML with fifteen schema types and two thousand words before any JavaScript runs. Having done that, the checklist had no further questions — while 97.8% of our 404 traffic came from things that will never send us a reader, the biggest uncached object on the domain was not a page, and the page-weight figure we were about to publish was measuring bytes nobody receives.

The checklist tells you how to be readable. The request log tells you who is reading. They are different instruments and only one of them is in the standard kit.

Sources and method

FigureSourceDate
1,797 404s in 48 hours; 40 from verified search crawlers; ~547 credential scannersCloudflare edge logs for our own domain2026-09-20 window
394 /alternatives/ redirects rejected after checking requester identitySame logs; our own decision record2026-09-20
218.3 KB uncompressed vs 23.7 KB brotli across 40 built tool pages; 9.2× ratioOur own build output, measured both ways2026-09-20
Hydration payload 58% uncompressed / 40% on the wire; 56% vs 33% savingSame measurement2026-09-20
Social preview image: 174 requests, 31.1 MB per 72hCloudflare edge logs2026-09-20
87.8% cache hit rate on tool pages; ~27,500 requests/day flat for 19 daysCloudflare analytics2026-09-20
Unresolved 1.5× gap; ~50% error bar on savings estimatesOur own diagnosis document, stated at the top of it2026-09-20
27 force-static route files; 7 with dynamicParams false; 1,956 words in raw HTML; 15 schema typesThis repository and its built output2026-10-07
~30% of sites with JS-dependent content; ~12% implementing schemaSemrush's trend brief — not verified by usas published
Semrush One and AI visibility wording, verbatimSemrush's own home page and AI visibility feature page2026-10-07
Site Audit crawling with JavaScript rendering enabledSemrush's Creator Accelerator briefretrieved 2026-09-22
"Try Semrush free for seven days. Cancel anytime"Semrush's pricing page2026-10-07
Screenshot wording "Win in search, era after era" (headline since changed)Semrush's home page, captured by us2026-09-26

Written by Sohail Akhtar. Figures describe one domain on the dates given and should be read as one domain's measurements, not as norms.

Explore More AI Tools

Browse our curated directory of 883+ verified AI tools.

Browse Directory

Some links may be affiliate links. We may earn a small commission at no extra cost to you.