Skip to content

Google Indexing Has Three Gates. We Measured Where 40 Random Pages Stopped.

Sohail Akhtar
Researched by Sohail AkhtarTheToolsVerse
October 10, 202611 min read
Cover image: of 40 random sitemap URLs inspected in Google Search Console on 7 October 2026, 26 had ever been crawled, 19 were indexed, and none received a Google click in 97 days.

Disclosure

This article is promotional content produced as part of the Semrush Creator Accelerator programme. It contains an affiliate link to Semrush — if you buy through it we may earn a commission at no extra cost to you — and we may also receive a reward from Semrush for publishing it. All measurements in this article are our own. Semrush did not supply, review or approve the data, and the conclusions we draw are ours.

Semrush's trend brief on Google indexing makes a claim that is easy to repeat and hard to check: that Google now applies its quality judgement "at the indexing layer, not just the ranking layer", and that pages can sit in "Crawled — currently not indexed" and "Discovered — currently not indexed" indefinitely as a result.

Check the gate you control first

Check whether Google can crawl and index your pages with Semrush One — start your 7-day free trial. Its Site Audit tests crawlability, indexability, robots directives and sitemap coverage: of the three gates below, the only one that is fully in your control. The other two are what this article is about.

We have a site that can check it. Our sitemap lists 612 URLs we ask Google to index. On 7 October 2026 we drew 40 of them at random, asked Google's URL Inspection API about each one, and then looked up what each had earned in Google Search between 1 July and 5 October.

Of 40 pages, 26 had ever been crawled. 19 were indexed. 10 had been shown to anyone. One had been shown ten times or more. None had been clicked.

Forty random sitemap URLs through Google's gates: 40 in the sitemap, 26 ever crawled, 19 indexed, 10 shown at least once, 1 shown ten or more times, 0 clicked. All 40 together drew 53 Google impressions in 97 days.
Forty random sitemap URLs through Google's gates: 40 in the sitemap, 26 ever crawled, 19 indexed, 10 shown at least once, 1 shown ten or more times, 0 clicked. All 40 together drew 53 Google impressions in 97 days.

This article follows those forty pages through each gate, puts the Bing numbers for the same pages beside them, and separates the part of the problem a site owner can fix from the part they cannot. One of the three findings we did not expect, and it is the most useful one.

Why this site is a clean test

Two things make our domain unusually readable for this question.

The same pages get a completely different verdict on another engine. Over the same quarter, the 477 tool pages in our sitemap drew 8,337 impressions and 6 clicks from Google, and 521,887 impressions and 3,367 clicks from Bing. Same HTML, same server, same Cloudflare configuration, same sitemap.

The same 477 tool pages: Bing 3,367 clicks and 521,887 impressions; Google 6 clicks and 8,337 impressions, July to early October 2026.
The same 477 tool pages: Bing 3,367 clicks and 521,887 impressions; Google 6 clicks and 8,337 impressions, July to early October 2026.

Semrush's brief describes exactly this check — practitioners "using Bing and DuckDuckGo indexing as a baseline to determine whether their Google problems are quality-based or technical." A gap of that size between two crawlers reading identical files cannot be a server problem. A shared cause cannot produce opposite outcomes on two consumers of the same output.

We already know when the threshold arrived. On 8 January 2026 Google showed our pages 13,448 times. On 9 January it showed them 486 times — a fall of 96% in one day. Average position did not collapse with it; it improved, from 16.0 to 12.7, and click-through rate rose from 0.40% to 12.87% within four days. That is not the shape of pages falling out of an index. It is the shape of an index that still contains you and has stopped showing you for anything but the queries where you are the obvious answer. Search Console's Manual Actions panel showed no issues when we checked it in July.

We are, in other words, a site the brief is describing. That makes us a poor example of how to succeed and a good instrument for measuring what the threshold looks like from inside.

How we drew the sample

  • Population: the 612 URLs in our live sitemap.xml on 7 October 2026 — every URL we are asking Google to index, nothing we have noindexed or redirected.
  • Draw: sorted alphabetically, shuffled with a seeded Fisher–Yates shuffle (seed 20261007), first 40 taken. The seed is printed so anyone with the file can redraw the same 40.
  • Index state: Google's URL Inspection API, one call per URL, 7 October 2026.
  • Exposure: Google Search Console page-level impressions and clicks, 1 July – 5 October 2026, www and apex merged.
  • Comparison: Bing Webmaster Tools page traffic for the same period, in weekly buckets from 3 July to 2 October.

Forty is a small sample, and every percentage below carries a wide interval. "19 of 40 indexed" is 47.5%, but the 95% confidence interval runs from 33% to 63%. We report counts first for that reason.

The census

Google's verdictPages
Submitted and indexed19
Discovered – currently not indexed11
Crawled – currently not indexed4
URL is unknown to Google3
Excluded by 'noindex' tag2
Blocked by robots.txt1

The last two rows should be impossible. Nothing in our sitemap is noindexed or blocked. We will come back to them, because they turned out to be the part we could actually fix.

Gate one: never fetched

Fourteen of forty pages — 35% — had never been crawled at all. Eleven sat in Discovered – currently not indexed with a last-crawl date of "Never"; three more were unknown to Google despite being in the sitemap.

This is the gate the brief is most right about. Google has these URLs. It has decided not to spend a fetch on them. That decision is made before Google has read a word of the page, so it cannot be a judgement on the page. It is a judgement on the site.

It is also not slow in a uniform way. On 24 September we published two articles on the same afternoon, in the same sitemap, linked from the same blog index. Thirteen days later one was Submitted and indexed, last crawled on 26 September. The other — our measurement of ChatGPT retrievals and citations — was still Discovered – currently not indexed, never fetched. Nothing about the two pages explains the difference from where we sit. That is what a budget looks like from the outside: something gets spent, and something else waits.

The only crawl figure we have for this is from Google Search Console's crawl stats in July 2026, when 8% of Googlebot's requests to our site were discovery and 92% were refreshes of pages it already knew. A site Google trusts less gets a smaller share of attention for new URLs, and new URLs are exactly where a growing site needs it.

Gate two: fetched, read, declined

Four pages had been crawled and refused — Crawled – currently not indexed. Google spent the fetch, read the page, and chose not to keep it.

One of the four is our own listing page for one of Semrush's AI tools, crawled on 3 September and declined. Bing showed the same page 20 times in the same quarter. Another, our RemNote listing, was declined by Google and shown 286 times by Bing.

This is the gate where the brief's advice lands hardest: "Stop treating 'Crawled — currently not indexed' as a bug to fix with sitemaps and URL inspection requests." We agree, with one qualification that the next two sections earn. Resubmitting a page Google has read and declined asks it to repeat a decision it has already made, on the same evidence.

Gate three: indexed, and shown to nobody

This is the gate that most indexing advice stops before.

Nineteen of the forty were indexed. Nine of those nineteen were not shown in Google Search once in 97 days. Ten were shown at least once. One was shown ten times or more. Together, all forty pages earned 53 Google impressions in the quarter and no clicks.

The clearest case is our Meta AI listing. Google's URL Inspection reports it as Submitted and indexed, last crawled 1 September 2026, with valid Breadcrumb and Review snippet markup. In the quarter it drew zero Google impressions. Bing showed it 1,094 times and sent 9 clicks.

The brief names this too — pages "marked 'Indexed' in Google Search Console that are completely invisible in live SERPs." It frames that as a reason to distrust Search Console. Our reading is narrower. Search Console is not wrong. The page is indexed. "Indexed" has simply stopped meaning "eligible to be shown", and treating index coverage as the success metric is the mistake.

Indexed is a state. Served is the outcome. The only report that measures the second is the Performance report, page by page, and almost nobody reads it against the coverage report.

The gate we built ourselves

Three of the forty showed states that our sitemap says cannot exist: two Excluded by 'noindex' tag and one Blocked by robots.txt.

We checked all three live. All three currently serve index, follow, and robots.txt blocks none of them. Then we checked the dates:

PageGoogle's last crawlWhat Google sawWhen we fixed it
a text-to-video tool listing22 June 2026noindexreleased 25 July
a media-monitoring tool listing8 July 2026noindexreleased 27 August
an alternatives page21 July 2026blocked by robots.txtrule removed 5 August

In every case Google's view of the page is older than our fix. We changed the page; Google has not been back. Two of these have now been carrying a fix for over two months that Google has never seen.

This qualifies the brief's advice in a way that matters. "Crawled – currently not indexed is a quality verdict, not a bug" is true of pages Google has read in their current form. It is not true of a page whose last crawl predates a technical change — and on a site where Google spends little attention on refreshes for low-value URLs, that window can last months. 7.5% of our random sample was in it, which on 612 URLs is roughly 46 pages Google is judging on a version that no longer exists — though with an interval from 3% to 20%, that estimate is soft.

There are two genuinely different problems wearing the same label in Search Console. One is a verdict. The other is stale evidence, and it is the one a site owner can act on.

What the forty do not show

  • They do not show why Google declines any particular page. URL Inspection reports a state, not a reason. Everything we say about cause is inference from the pattern, and the January signature, not from any statement Google made about our site.
  • Forty is a sample, not a census. The intervals are wide. The direction of the findings is robust; the decimals are not.
  • The Bing comparison is about totals, not coverage. Bing's page export lists top pages only, so "Bing showed page X zero times" can mean we lack the row, not that Bing ignored it. We only compare totals, and individual pages Bing did report.
  • They do not show that pruning or rewriting would fix anything. We have done both, at scale, earlier this year. The January threshold has not lifted. We are reporting an observation, not a recovery.
  • One more thing pruning cost us, which we have written about separately: pages removed to satisfy Google's threshold included pages ChatGPT was citing. The two systems want different things.

Run this on your own site

The whole method needs a sitemap, Search Console, and an afternoon.

  1. Sample from your sitemap, not your analytics. The sitemap is the list of pages you are asking Google to index; analytics only shows pages that already won. Use a printed seed.
  2. Inspect every sampled URL and record the state and the last-crawl date. The date is the column that separates a verdict from stale evidence.
  3. Join each URL to Performance data — impressions and clicks for the last quarter. Count how many indexed pages were never shown. That number, not index coverage, is the one to track.
  4. For every non-indexed page, compare the last-crawl date with the date you last changed it. If the crawl is older, Google is judging a page that no longer exists. That is the one case where requesting a recrawl is worth doing.
  5. Check the same pages on Bing Webmaster Tools. If Bing serves them and Google does not, stop looking for a robots or server fault. You will not find one.
  6. Audit for the barriers you built yourself — stray noindex, robots rules, sitemap entries that conflict with either, orphaned pages. This is the only gate that is fully in your control, and a regular crawler catches it faster than a quarterly look at Search Console.

Steps one to five are free. Step six is where a crawler earns its keep.

Where Semrush fits — and where it does not

Of the three gates, a site audit tool can see one: the barriers on your own side. Semrush's Site Audit page — which now doubles as a free SEO checker — says, checked 7 October 2026, that it tests "crawlability, indexability, status codes, robots directives, sitemap coverage, canonical tags, and structured data", and describes the failure we found directly — "Robots.txt blocks, noindex tags, and orphan pages stop search engines from seeing your content. We list every barrier so you can remove it." It runs "140+ technical and on-page checks", and it lists AI Search Health monitoring to "Spot issues that can reduce your visibility in LLMs and AI-driven results."

Semrush's own illustration of Site Audit's issues list, from its Site Audit feature page (retrieved 9 October 2026): errors, warnings and notices filtered into categories including Crawlability and AI Search, with an "About the issue / How to fix" panel. Semrush's illustration, not our site.
Semrush's own illustration of Site Audit's issues list, from its Site Audit feature page (retrieved 9 October 2026): errors, warnings and notices filtered into categories including Crawlability and AI Search, with an "About the issue / How to fix" panel. Semrush's illustration, not our site.

That is the gate where our three stale pages began. A crawler that had flagged a noindex on a page we were simultaneously listing in the sitemap would have caught the conflict on the day it existed, rather than leaving us to discover, three months later, that Google was still looking at it.

The four pages Google read and declined at gate two raise a different question — not whether Google can reach the page, but how the page compares with what Google does show. That is what On Page SEO Checker, the other On-Page & Technical SEO tool in the SEO Toolkit, is for. Its Knowledge Base entry, checked 9 October 2026, says it "offers a complete and structured list of things you could do to improve the ranks of pages on your website", including "on-page SEO ideas, semantically related words to include on your pages, target content length and readability, and backlink prospects." It cannot tell you why Google declined a page. It can show where a declined page falls short of the pages that rank — which is the only part of a quality verdict you can act on.

What no third-party crawler can see is Google's decision. Whether a fetched page is kept, and whether a kept page is shown, lives in Google's systems; only Search Console reports it. Treat the two as complementary: an audit for the barriers you can remove, URL Inspection and the Performance report for the verdict you cannot.

Semrush One bundles Site Audit with the rest of the SEO Toolkit and the AI Visibility Toolkit. Semrush describes it as letting you "measure your visibility across SERPs and LLMs, spot competitive gaps, and audit your site’s AI readiness" — which is the cross-engine view this article had to assemble by hand from two webmaster consoles.

Position Tracking, also part of Semrush One, covers a blind spot this site has already fallen into. When Google stopped serving us on 9 January, Search Console's average position improved, from 16.0 to 12.7, on the same day impressions fell 96%. An average over whatever queries survive can only tell you about the queries that survived. Semrush's Position Tracking page draws the same line, checked 9 October 2026: "GSC shows you aggregate data with a delay and no competitor visibility. Position Tracking gives you daily updates, keyword-level granularity, mobile and desktop splits, ZIP-code-level local data, and side-by-side competitor rankings", with "Real-time alerts and notifications". A daily check on a fixed set of long-tail keywords is built to show the shape of an event like ours, those keywords dropping out, on the day it happens.

Semrush's own illustration of a Position Tracking alert, from its Position Tracking feature page (retrieved 9 October 2026): "Notify me when positions for any keyword for domain.com enter the top 100", with a "Set up alert" button. Semrush's illustration, not our account.
Semrush's own illustration of a Position Tracking alert, from its Position Tracking feature page (retrieved 9 October 2026): "Notify me when positions for any keyword for domain.com enter the top 100", with a "Set up alert" button. Semrush's illustration, not our account.

If you want to try the audit side before paying for it, Semrush's checker page states "You can run one check per day for free". For the full toolkit there is a 7-day free trial: Semrush's pricing page reads "Try Semrush free for seven days. Cancel anytime", and its Semrush One page offers "Try free for 7 days" with "Unlimited access to all Semrush One tools" — both checked 7 October 2026, the trial for new users only. Read those as Semrush's terms rather than our test; the measurements in this article were made with Google's and Bing's own consoles.

Semrush One's own product page: "Try free for 7 days" button with "Unlimited access to all Semrush One tools" beneath it. Captured from Semrush's page on 26 September 2026.
Semrush One's own product page: "Try free for 7 days" button with "Unlimited access to all Semrush One tools" beneath it. Captured from Semrush's page on 26 September 2026.

Semrush One — 7-day free trial

Three of our forty pages were being judged on a version we had already fixed. Site Audit catches the barriers you build yourself on the day they exist; Position Tracking shows a drop on the day it happens. Search Console told us months later.

Start your 7-day free trial

What we actually learned

The brief is right that Google's threshold now operates before ranking. On our site it operates in three places, and they need three different responses.

At the first gate, Google does not fetch. Fourteen of forty pages had never been crawled. Nothing on those pages can be the reason, because nothing on them has been read.

At the third gate, Google indexes and does not show. Nine of nineteen indexed pages drew no impressions in a quarter, including one that Bing showed over a thousand times. "Indexed" is not the finish line, and coverage is the wrong number to celebrate.

And between them sits a gate we built ourselves and forgot about: three pages Google was still judging on a version we had changed months earlier. That is the only part of this that a site owner can fix by Monday — and the only case where asking Google to look again is not a waste of everyone's time.

Fix that gate first. Check whether Google can crawl and index your pages with Semrush One — start your 7-day free trial — and keep Search Console open beside it for the verdicts no outside tool can see.

Try Semrush One free for 7 days

Site Audit for the crawl and index barriers on your own side, Position Tracking for drops on the day they happen. Free for 7 days for new users.

Start your 7-day free trial

Sources and method

FigureSourceDate
612 URLs in the sitemap; 40 drawn with seed 20261007Our live sitemap.xml; draw reproducible from the saved file2026-10-07
Index state and last-crawl date for each of the 40Google Search Console URL Inspection API, our verified property2026-10-07
53 impressions, 0 clicks across the 40; 9 of 19 indexed pages never shownGoogle Search Console Performance, page level, www and apex merged2026-07-01 → 2026-10-05
/tools/* pages: Google 8,337 impressions / 6 clicks; Bing 521,887 / 3,367Search Console as above; Bing Webmaster Tools page traffic, weekly bucketsGoogle 07-01 → 10-05; Bing weeks of 07-03 → 10-02
Meta AI listing: Bing 1,094 impressions / 9 clicks; RemNote 286; AI Visibility Toolkit listing 20Bing Webmaster Tools, same exportas above
Two articles of 24 September: one indexed, one never crawledURL Inspection API2026-10-07
Live index, follow on the three stale pages; fix datesLive HTML fetched by us; our own commit history2026-10-07
13,448 → 486 impressions; position 16.0 → 12.7; CTR 0.40% → 12.87%Google Search Console Performance, daily2026-01-08 → 2026-01-13
8% discovery / 92% refresh crawl requests; no manual actionsGoogle Search Console crawl stats and Manual Actions panelJuly 2026
Quotations from the indexing trend briefSemrush Creator Accelerator, Content & Promo Ideaspage edited 2026-10-05
Site Audit and Semrush One wording; free check; 7-day trialSemrush's own Site Audit, Semrush One and pricing pages2026-10-07
On Page SEO Checker descriptionSemrush Knowledge Base: On Page SEO Checker2026-10-09
Position Tracking vs GSC; "Real-time alerts and notifications"Semrush's Position Tracking feature page2026-10-09
Site Audit issues-list and Position Tracking alert screenshots (Semrush's illustrations)Semrush's Site Audit and Position Tracking feature pagesretrieved 2026-10-09

Written by Sohail Akhtar. Every figure here is a measurement on a stated date and can be re-derived from the saved exports.

Explore More AI Tools

Browse our curated directory of 883+ verified AI tools.

Browse Directory

Some links may be affiliate links. We may earn a small commission at no extra cost to you.