Google Indexing Has Three Gates. We Measured Where 40 Random Pages Stopped.
Disclosure
This article is promotional content produced as part of the Semrush Creator Accelerator programme. It contains an affiliate link to Semrush — if you buy through it we may earn a commission at no extra cost to you — and we may also receive a reward from Semrush for publishing it. All measurements in this article are our own. Semrush did not supply, review or approve the data, and the conclusions we draw are ours.
Semrush's trend brief on Google indexing makes a claim that is easy to repeat and hard to check: that Google now applies its quality judgement "at the indexing layer, not just the ranking layer", and that pages can sit in "Crawled — currently not indexed" and "Discovered — currently not indexed" indefinitely as a result.
Check the gate you control first
Check whether Google can crawl and index your pages with Semrush One — start your 7-day free trial. Its Site Audit tests crawlability, indexability, robots directives and sitemap coverage: of the three gates below, the only one that is fully in your control. The other two are what this article is about.
We have a site that can check it. Our sitemap lists 612 URLs we ask Google to index. On 7 October 2026 we drew 40 of them at random, asked Google's URL Inspection API about each one, and then looked up what each had earned in Google Search between 1 July and 5 October.
Of 40 pages, 26 had ever been crawled. 19 were indexed. 10 had been shown to anyone. One had been shown ten times or more. None had been clicked.
This article follows those forty pages through each gate, puts the Bing numbers for the same pages beside them, and separates the part of the problem a site owner can fix from the part they cannot. One of the three findings we did not expect, and it is the most useful one.
Why this site is a clean test
Two things make our domain unusually readable for this question.
The same pages get a completely different verdict on another engine. Over the same quarter, the 477 tool pages in our sitemap drew 8,337 impressions and 6 clicks from Google, and 521,887 impressions and 3,367 clicks from Bing. Same HTML, same server, same Cloudflare configuration, same sitemap.
Semrush's brief describes exactly this check — practitioners "using Bing and DuckDuckGo indexing as a baseline to determine whether their Google problems are quality-based or technical." A gap of that size between two crawlers reading identical files cannot be a server problem. A shared cause cannot produce opposite outcomes on two consumers of the same output.
We already know when the threshold arrived. On 8 January 2026 Google showed our pages 13,448 times. On 9 January it showed them 486 times — a fall of 96% in one day. Average position did not collapse with it; it improved, from 16.0 to 12.7, and click-through rate rose from 0.40% to 12.87% within four days. That is not the shape of pages falling out of an index. It is the shape of an index that still contains you and has stopped showing you for anything but the queries where you are the obvious answer. Search Console's Manual Actions panel showed no issues when we checked it in July.
We are, in other words, a site the brief is describing. That makes us a poor example of how to succeed and a good instrument for measuring what the threshold looks like from inside.
How we drew the sample
- Population: the 612 URLs in our live
sitemap.xmlon 7 October 2026 — every URL we are asking Google to index, nothing we have noindexed or redirected. - Draw: sorted alphabetically, shuffled with a seeded Fisher–Yates shuffle (seed
20261007), first 40 taken. The seed is printed so anyone with the file can redraw the same 40. - Index state: Google's URL Inspection API, one call per URL, 7 October 2026.
- Exposure: Google Search Console page-level impressions and clicks, 1 July – 5 October 2026,
wwwand apex merged. - Comparison: Bing Webmaster Tools page traffic for the same period, in weekly buckets from 3 July to 2 October.
Forty is a small sample, and every percentage below carries a wide interval. "19 of 40 indexed" is 47.5%, but the 95% confidence interval runs from 33% to 63%. We report counts first for that reason.
The census
| Google's verdict | Pages |
|---|---|
| Submitted and indexed | 19 |
| Discovered – currently not indexed | 11 |
| Crawled – currently not indexed | 4 |
| URL is unknown to Google | 3 |
| Excluded by 'noindex' tag | 2 |
| Blocked by robots.txt | 1 |
The last two rows should be impossible. Nothing in our sitemap is noindexed or blocked. We will come back to them, because they turned out to be the part we could actually fix.
Gate one: never fetched
Fourteen of forty pages — 35% — had never been crawled at all. Eleven sat in Discovered – currently not indexed with a last-crawl date of "Never"; three more were unknown to Google despite being in the sitemap.
This is the gate the brief is most right about. Google has these URLs. It has decided not to spend a fetch on them. That decision is made before Google has read a word of the page, so it cannot be a judgement on the page. It is a judgement on the site.
It is also not slow in a uniform way. On 24 September we published two articles on the same afternoon, in the same sitemap, linked from the same blog index. Thirteen days later one was Submitted and indexed, last crawled on 26 September. The other — our measurement of ChatGPT retrievals and citations — was still Discovered – currently not indexed, never fetched. Nothing about the two pages explains the difference from where we sit. That is what a budget looks like from the outside: something gets spent, and something else waits.
The only crawl figure we have for this is from Google Search Console's crawl stats in July 2026, when 8% of Googlebot's requests to our site were discovery and 92% were refreshes of pages it already knew. A site Google trusts less gets a smaller share of attention for new URLs, and new URLs are exactly where a growing site needs it.
Gate two: fetched, read, declined
Four pages had been crawled and refused — Crawled – currently not indexed. Google spent the fetch, read the page, and chose not to keep it.
One of the four is our own listing page for one of Semrush's AI tools, crawled on 3 September and declined. Bing showed the same page 20 times in the same quarter. Another, our RemNote listing, was declined by Google and shown 286 times by Bing.
This is the gate where the brief's advice lands hardest: "Stop treating 'Crawled — currently not indexed' as a bug to fix with sitemaps and URL inspection requests." We agree, with one qualification that the next two sections earn. Resubmitting a page Google has read and declined asks it to repeat a decision it has already made, on the same evidence.
Gate three: indexed, and shown to nobody
This is the gate that most indexing advice stops before.
Nineteen of the forty were indexed. Nine of those nineteen were not shown in Google Search once in 97 days. Ten were shown at least once. One was shown ten times or more. Together, all forty pages earned 53 Google impressions in the quarter and no clicks.
The clearest case is our Meta AI listing. Google's URL Inspection reports it as Submitted and indexed, last crawled 1 September 2026, with valid Breadcrumb and Review snippet markup. In the quarter it drew zero Google impressions. Bing showed it 1,094 times and sent 9 clicks.
The brief names this too — pages "marked 'Indexed' in Google Search Console that are completely invisible in live SERPs." It frames that as a reason to distrust Search Console. Our reading is narrower. Search Console is not wrong. The page is indexed. "Indexed" has simply stopped meaning "eligible to be shown", and treating index coverage as the success metric is the mistake.
Indexed is a state. Served is the outcome. The only report that measures the second is the Performance report, page by page, and almost nobody reads it against the coverage report.
The gate we built ourselves
Three of the forty showed states that our sitemap says cannot exist: two Excluded by 'noindex' tag and one Blocked by robots.txt.
We checked all three live. All three currently serve index, follow, and robots.txt blocks none of them. Then we checked the dates:
| Page | Google's last crawl | What Google saw | When we fixed it |
|---|---|---|---|
| a text-to-video tool listing | 22 June 2026 | noindex | released 25 July |
| a media-monitoring tool listing | 8 July 2026 | noindex | released 27 August |
| an alternatives page | 21 July 2026 | blocked by robots.txt | rule removed 5 August |
In every case Google's view of the page is older than our fix. We changed the page; Google has not been back. Two of these have now been carrying a fix for over two months that Google has never seen.
This qualifies the brief's advice in a way that matters. "Crawled – currently not indexed is a quality verdict, not a bug" is true of pages Google has read in their current form. It is not true of a page whose last crawl predates a technical change — and on a site where Google spends little attention on refreshes for low-value URLs, that window can last months. 7.5% of our random sample was in it, which on 612 URLs is roughly 46 pages Google is judging on a version that no longer exists — though with an interval from 3% to 20%, that estimate is soft.
There are two genuinely different problems wearing the same label in Search Console. One is a verdict. The other is stale evidence, and it is the one a site owner can act on.
What the forty do not show
- They do not show why Google declines any particular page. URL Inspection reports a state, not a reason. Everything we say about cause is inference from the pattern, and the January signature, not from any statement Google made about our site.
- Forty is a sample, not a census. The intervals are wide. The direction of the findings is robust; the decimals are not.
- The Bing comparison is about totals, not coverage. Bing's page export lists top pages only, so "Bing showed page X zero times" can mean we lack the row, not that Bing ignored it. We only compare totals, and individual pages Bing did report.
- They do not show that pruning or rewriting would fix anything. We have done both, at scale, earlier this year. The January threshold has not lifted. We are reporting an observation, not a recovery.
- One more thing pruning cost us, which we have written about separately: pages removed to satisfy Google's threshold included pages ChatGPT was citing. The two systems want different things.
Run this on your own site
The whole method needs a sitemap, Search Console, and an afternoon.
- Sample from your sitemap, not your analytics. The sitemap is the list of pages you are asking Google to index; analytics only shows pages that already won. Use a printed seed.
- Inspect every sampled URL and record the state and the last-crawl date. The date is the column that separates a verdict from stale evidence.
- Join each URL to Performance data — impressions and clicks for the last quarter. Count how many indexed pages were never shown. That number, not index coverage, is the one to track.
- For every non-indexed page, compare the last-crawl date with the date you last changed it. If the crawl is older, Google is judging a page that no longer exists. That is the one case where requesting a recrawl is worth doing.
- Check the same pages on Bing Webmaster Tools. If Bing serves them and Google does not, stop looking for a robots or server fault. You will not find one.
- Audit for the barriers you built yourself — stray
noindex, robots rules, sitemap entries that conflict with either, orphaned pages. This is the only gate that is fully in your control, and a regular crawler catches it faster than a quarterly look at Search Console.
Steps one to five are free. Step six is where a crawler earns its keep.
Where Semrush fits — and where it does not
Of the three gates, a site audit tool can see one: the barriers on your own side. Semrush's Site Audit page — which now doubles as a free SEO checker — says, checked 7 October 2026, that it tests "crawlability, indexability, status codes, robots directives, sitemap coverage, canonical tags, and structured data", and describes the failure we found directly — "Robots.txt blocks, noindex tags, and orphan pages stop search engines from seeing your content. We list every barrier so you can remove it." It runs "140+ technical and on-page checks", and it lists AI Search Health monitoring to "Spot issues that can reduce your visibility in LLMs and AI-driven results."

That is the gate where our three stale pages began. A crawler that had flagged a noindex on a page we were simultaneously listing in the sitemap would have caught the conflict on the day it existed, rather than leaving us to discover, three months later, that Google was still looking at it.
The four pages Google read and declined at gate two raise a different question — not whether Google can reach the page, but how the page compares with what Google does show. That is what On Page SEO Checker, the other On-Page & Technical SEO tool in the SEO Toolkit, is for. Its Knowledge Base entry, checked 9 October 2026, says it "offers a complete and structured list of things you could do to improve the ranks of pages on your website", including "on-page SEO ideas, semantically related words to include on your pages, target content length and readability, and backlink prospects." It cannot tell you why Google declined a page. It can show where a declined page falls short of the pages that rank — which is the only part of a quality verdict you can act on.
What no third-party crawler can see is Google's decision. Whether a fetched page is kept, and whether a kept page is shown, lives in Google's systems; only Search Console reports it. Treat the two as complementary: an audit for the barriers you can remove, URL Inspection and the Performance report for the verdict you cannot.
Semrush One bundles Site Audit with the rest of the SEO Toolkit and the AI Visibility Toolkit. Semrush describes it as letting you "measure your visibility across SERPs and LLMs, spot competitive gaps, and audit your site’s AI readiness" — which is the cross-engine view this article had to assemble by hand from two webmaster consoles.
Position Tracking, also part of Semrush One, covers a blind spot this site has already fallen into. When Google stopped serving us on 9 January, Search Console's average position improved, from 16.0 to 12.7, on the same day impressions fell 96%. An average over whatever queries survive can only tell you about the queries that survived. Semrush's Position Tracking page draws the same line, checked 9 October 2026: "GSC shows you aggregate data with a delay and no competitor visibility. Position Tracking gives you daily updates, keyword-level granularity, mobile and desktop splits, ZIP-code-level local data, and side-by-side competitor rankings", with "Real-time alerts and notifications". A daily check on a fixed set of long-tail keywords is built to show the shape of an event like ours, those keywords dropping out, on the day it happens.

If you want to try the audit side before paying for it, Semrush's checker page states "You can run one check per day for free". For the full toolkit there is a 7-day free trial: Semrush's pricing page reads "Try Semrush free for seven days. Cancel anytime", and its Semrush One page offers "Try free for 7 days" with "Unlimited access to all Semrush One tools" — both checked 7 October 2026, the trial for new users only. Read those as Semrush's terms rather than our test; the measurements in this article were made with Google's and Bing's own consoles.

Semrush One — 7-day free trial
Three of our forty pages were being judged on a version we had already fixed. Site Audit catches the barriers you build yourself on the day they exist; Position Tracking shows a drop on the day it happens. Search Console told us months later.
Start your 7-day free trialWhat we actually learned
The brief is right that Google's threshold now operates before ranking. On our site it operates in three places, and they need three different responses.
At the first gate, Google does not fetch. Fourteen of forty pages had never been crawled. Nothing on those pages can be the reason, because nothing on them has been read.
At the third gate, Google indexes and does not show. Nine of nineteen indexed pages drew no impressions in a quarter, including one that Bing showed over a thousand times. "Indexed" is not the finish line, and coverage is the wrong number to celebrate.
And between them sits a gate we built ourselves and forgot about: three pages Google was still judging on a version we had changed months earlier. That is the only part of this that a site owner can fix by Monday — and the only case where asking Google to look again is not a waste of everyone's time.
Fix that gate first. Check whether Google can crawl and index your pages with Semrush One — start your 7-day free trial — and keep Search Console open beside it for the verdicts no outside tool can see.
Try Semrush One free for 7 days
Site Audit for the crawl and index barriers on your own side, Position Tracking for drops on the day they happen. Free for 7 days for new users.
Start your 7-day free trialSources and method
| Figure | Source | Date |
|---|---|---|
| 612 URLs in the sitemap; 40 drawn with seed 20261007 | Our live sitemap.xml; draw reproducible from the saved file | 2026-10-07 |
| Index state and last-crawl date for each of the 40 | Google Search Console URL Inspection API, our verified property | 2026-10-07 |
| 53 impressions, 0 clicks across the 40; 9 of 19 indexed pages never shown | Google Search Console Performance, page level, www and apex merged | 2026-07-01 → 2026-10-05 |
| /tools/* pages: Google 8,337 impressions / 6 clicks; Bing 521,887 / 3,367 | Search Console as above; Bing Webmaster Tools page traffic, weekly buckets | Google 07-01 → 10-05; Bing weeks of 07-03 → 10-02 |
| Meta AI listing: Bing 1,094 impressions / 9 clicks; RemNote 286; AI Visibility Toolkit listing 20 | Bing Webmaster Tools, same export | as above |
| Two articles of 24 September: one indexed, one never crawled | URL Inspection API | 2026-10-07 |
Live index, follow on the three stale pages; fix dates | Live HTML fetched by us; our own commit history | 2026-10-07 |
| 13,448 → 486 impressions; position 16.0 → 12.7; CTR 0.40% → 12.87% | Google Search Console Performance, daily | 2026-01-08 → 2026-01-13 |
| 8% discovery / 92% refresh crawl requests; no manual actions | Google Search Console crawl stats and Manual Actions panel | July 2026 |
| Quotations from the indexing trend brief | Semrush Creator Accelerator, Content & Promo Ideas | page edited 2026-10-05 |
| Site Audit and Semrush One wording; free check; 7-day trial | Semrush's own Site Audit, Semrush One and pricing pages | 2026-10-07 |
| On Page SEO Checker description | Semrush Knowledge Base: On Page SEO Checker | 2026-10-09 |
| Position Tracking vs GSC; "Real-time alerts and notifications" | Semrush's Position Tracking feature page | 2026-10-09 |
| Site Audit issues-list and Position Tracking alert screenshots (Semrush's illustrations) | Semrush's Site Audit and Position Tracking feature pages | retrieved 2026-10-09 |
Written by Sohail Akhtar. Every figure here is a measurement on a stated date and can be re-derived from the saved exports.
Some links may be affiliate links. We may earn a small commission at no extra cost to you.
Related Articles
Explore more guides and reviews from our experts.