Retrieved 44 Times, Cited 13: Four Kinds of AI Visibility
Disclosure
This article is promotional content produced as part of the Semrush Creator Accelerator programme. It contains an affiliate link to Semrush — if you buy through it we may earn a commission at no extra cost to you — and we may also receive a reward from Semrush for publishing it. All measurements in this article are our own. Semrush did not supply, review or approve the data, and the conclusions we draw about their own trend brief are ours.
Four things get called "AI visibility". On one domain — ours — measured across one search console and one third-party corpus, they came back like this:
- Page-one Bing positions across a family of commercial queries, window 2026-04-10 to 09-18
- 44 appearances in ChatGPT's search results, measured 2026-09-24
- 13 source citations in ChatGPT, same call, same date
- Zero mentions in Google's AI Overview corpus, same date, broadest available scope
Four instruments, four numbers, one website. They do not agree with each other, and the gaps between them are the entire point — because until recently we were quoting one of them under the wrong name.
Four different things are being called "AI visibility"
Before any number is worth reading, the categories have to be separate. Here is how we separate them. This is our taxonomy, not an industry standard — but the distinctions are real and they are measurable.
1. Search-engine visibility. Whether a search engine ranks your pages, measured by that engine's own console. Impressions, clicks, average position. Oldest and least ambiguous of the four.
2. Retrieval. Whether an LLM's search step returns your page as a candidate when it goes looking. You were fetched. You were in the running.
3. Citation. Whether the model's answer actually names you as a source. You were fetched and used.
4. AI Overview visibility. Whether Google's own AI answers name you — a different surface, a different index, and no reason to assume it tracks any of the above.
These move independently. A number that does not say which of the four it is measuring is not a measurement, it is a vibe — and conflating the second with the third is an error we made ourselves and published. That comes later.
What Semrush's briefs say — both of them
This article responds to a Semrush trend brief titled "Bing Is ChatGPT's Gatekeeper: Visibility Starts Outside Google" (21 April 2026). Its claim, attributed in the brief to an analysis by LLMVLab, is that "~87% of ChatGPT citations match Bing top 20 results", and that Bing indexing is therefore a direct requirement for AI visibility.
We are, almost exactly, the site that brief describes — visibility starting outside Google — so we ran the measurements on ourselves. We have not seen LLMVLab's analysis and cannot verify the 87%; it stays theirs.
A second Semrush brief, four months later (31 August 2026), describes different machinery. It reports that since GPT-5.6 on 8 August 2026, ChatGPT "stopped 'searching the web and citing what ranks' and started pre-selecting a shortlist of trusted domains, then searching only those" — a trust filter running before retrieval, rather than retrieval following Bing's ranking.
Two accounts, four months apart, from one publisher, of a system neither of us can see inside. That is not a contradiction to score points off. It is a fast-moving field being reported twice. We have verified neither mechanism, and this article settles neither.
Why this site is an unusually clean case
Most sites rank on Google and Bing simultaneously, so their search visibility and their AI visibility are hopelessly confounded — you cannot tell which engine is doing the work.
Ours are separated, which is rarer and more useful.
On Bing we hold page-one positions across a family of commercial queries. On Google, measured 2026-09-24, we had 47 ranked keywords, none in the top ten, at an average position of 67.3 — our best being position 18. Not invisible; page two at best, and mostly page five and beyond. Five weeks earlier the same check returned 82 ranked keywords at an average position of 61.1, so the trend over that window is downward on both counts. Either way, Google traffic is not what is putting us in front of anyone.
The Semrush brief imagines a site that is strong on Google and absent from AI answers. We are the inverse of that, which makes the four measurements worth taking.
Measurement 1: Bing, from Bing
Source: Bing Webmaster Tools, our own verified property. Window: 10 April to 18 September 2026 — 24 weekly buckets, re-pulled 2026-09-24.
| Measure | Value |
|---|---|
| Impressions, whole window | 642,539 |
| Clicks, whole window | 5,384 |
| Site CTR | 0.84% |
| Impression-weighted average position | 5.60 |
That 642,539 is a five-month total, not a monthly figure, and it would be a much better-looking number if we let you assume otherwise. For a single named month: August 2026 — four weekly buckets — 169,931 impressions and 1,500 clicks.
The /free collection is the strongest part of it: across 25 rows, 59,896 impressions, 1,406 clicks, a 2.35% click-through rate — roughly 2.8× the site-wide average — at an average position of 4.47. Our single largest query overall is narrower and more branded: canva ai, at 175,964 impressions and 802 clicks, average position 4.9.
Two notes on method, because they change what these numbers mean.
Positions are impression-weighted, not an average of averages — which would let a page with a handful of impressions at a flattering rank pull the whole figure toward a position we do not hold.
Every position here is an average across 24 weeks, not a live rank. A query sitting at 2.1 on 200 impressions is a smaller claim than one at 4.5 on 10,677.
Why we threw away a better-looking number
We used to cite a scraped result putting us at #10 on Bing for best free ai tools. We have retired it, and the reason is the most transferable thing in this section.
Scraped SERP data, in our own logs, has been wrong about us in both directions. In one documented case a scrape reported that we did not rank for a query at all while Bing Webmaster showed 49 impressions at an average position of 6.1 for that same query. In one 2026-08-19 research pass, roughly half the Bing SERP responses came back for a different query than the one sent.
So where first-party console data and a scrape both exist, we use the console. For best free ai tools, Bing's own numbers are 168 impressions and an average position of 7.27 across the window. Less flattering than #10, and ours.
One limitation you cannot check
This is private console data. You cannot reproduce it. We can re-run it on demand, for free, and the method above is stated precisely so you can run the same measurement on your own property.
Measurements 2 and 3: retrieved 44 times, cited 13
Source: DataForSEO's LLM Mentions API, target_metrics/live. Measured 2026-09-24. Platform chat_gpt, location_code 2840, language_code en, subdomains included. $0.101 per call.
The API separates two things that ordinary language does not. search_scope: ["sources"] returns domains "cited as sources in LLM responses" — the citation definition. search_scope: ["any"] returns everything, including domains that merely turned up in the search results behind a response.
Two calls, same target, same minute:
| Scope | What it counts | Result |
|---|---|---|
sources | Cited as a source | 13 |
any | Every appearance | 56 |
Inside that any run, our domain appears in the search-results breakdown 25 times on the www host and 19 times on the bare host — the API returns them as separate keys, so they have to be added. 44 search-result appearances against 13 source citations.
In DataForSEO's collected corpus of ChatGPT responses, on that date, our domain was retrieved roughly 3.4 times more often than it was cited.
That sentence carries its own qualifiers deliberately, because it is the one most likely to be repeated without them. It describes one domain, on one date, inside one vendor's corpus. It is arithmetic on our own two numbers, not a law about ChatGPT.
We are not the first to notice this gap, and it would be dishonest to present it as a discovery. It has been measured at far greater scale than ours: an AirOps analysis reported by Search Engine Land examined 548,534 pages ChatGPT retrieved across 15,000 prompts and found that roughly 85% of retrieved pages were never cited. We have not reproduced that study and it stays theirs.
What one domain adds is a different kind of evidence: a specific, fully parameterised measurement you can place against the aggregate. Ours works out at roughly 70% retrieved-but-not-cited — the same phenomenon, somewhat less severe than the large-sample average. That is a single data point and we would not read a trend into it. But it is the difference between knowing a pattern exists and knowing where your own site sits inside it, which is the only version of this number that changes what you do on Monday.
The limit worth printing rather than burying: this is not a live query we issued to ChatGPT. It is DataForSEO's own collected corpus of LLM responses, and its size and composition are not disclosed to us. "13 citations" means thirteen in their corpus on that date. Any sentence that drops that qualifier is overstating what we know.
One more thing fell out of the same call. The largest source domain co-occurring with our brand in that mention set is Reddit, at 19 of 56 — more than either of our own hosts. That is co-occurrence inside our slice of the corpus, not a ranking of Reddit against us. It does sit alongside the later Semrush brief's observation that Reddit appears heavily in ChatGPT's retrieval set "without earning a single citation" — the same retrieval-versus-citation gap, reported from their side, measured from ours.
The number we got wrong
In an internal report dated 2026-08-19 we recorded "ChatGPT citations: 62", with an ai_search_volume of 690.
It was not a citation count. Re-measuring on three scopes on the same day gave sources 13 against an ai_search_volume of 146, and any 56 against 637. 62/690 matches the shape of the any run, not the sources run. The historical figure was almost certainly every appearance of any kind, filed under the word "citations" — off by roughly four times against the metric its own label described.
Nobody misled us. DataForSEO documents its scopes clearly. We recorded a number without recording which scope produced it, and the label did the rest. It then survived two internal review passes, because a number with a plausible name does not look broken. Only actually re-running the measurement caught it.
That is the practical warning in this article: a "mentions" figure whose scope you cannot state is not a citation count, and the name attached to it will not tell you.
Measurement 4: zero, and what a zero is worth
The same endpoint, the same domain, the same date, with platform: google — Google's AI Overview — returned zero.
We ran it on search_scope: ["any"], the broadest setting available, specifically so the result could not be dismissed as a narrow filter. Every aggregated array came back empty and no total object was returned at all, which is what a genuine zero looks like in this API and is distinguishable from a small number that has been filtered down.
The contrast sits inside one session: same endpoint, same domain, same minute, one platform returning 56 and the other returning nothing. It also reproduces. We have now run this check three times across five weeks — 2026-08-19, 2026-09-22 and 2026-09-24 — and it has returned zero every time.
What that zero is: an absence in one vendor's Google AI Overview corpus, in US/English, on one date. What it is not: evidence that Google has excluded us, penalised us, or will never cite us. Absence of evidence in a corpus of undisclosed size is a much smaller statement than it looks, and it is worth being precise about, because a zero is the easiest number in the world to over-read.
What this does and does not show
One domain. One date. One vendor's corpus. One search console. This is an n-of-1 measurement with a fully stated method, not a study. Its value is that the configuration is rare, not that it is representative. (Stating the method and its limits up front is how we handle everything we publish; we have also written separately about auditing our own accuracy, where the same discipline caught a different mistake.)
We did not test causation, and nothing here establishes it. Bing visibility and ChatGPT citation are two observations on the same domain. There is no control, no second domain, no intervention — we changed nothing and watched what happened to nothing. Any account of why one follows the other would be a story told over two data points.
What our numbers are consistent with is a selection step sitting between retrieval and citation: we are clearly retrievable, and cited a good deal less often than we are retrieved. That is the direction the later Semrush brief describes. Consistent with is as far as it goes. It does not demonstrate that mechanism, and it says nothing at all about the 87% figure in the earlier one.
So: whether Bing indexing gates the candidate pool, whether a trust filter gates the shortlist, or whether both are happening, is not settled here and we are not going to pretend otherwise. Two briefs from one publisher describe two mechanisms four months apart. We have one site's numbers. Those are different kinds of object.
And the numbers will move. Between our two measurements five weeks apart, the any-scope figure fell by nearly a tenth — the corpus changes underneath you while you are not looking. Which is exactly why the useful part of this article is not the four numbers. It is that they are four numbers rather than one, that each carries the scope and date that make it meaningful, and that a method you can re-run beats a statistic you have memorised.
Measure your own four numbers
- Verify a Bing Webmaster property and read the query stats. Free, and most teams have never done it — which means one of these four dimensions is simply invisible to them.
- Weight positions by impressions. Averaging averages flatters low-traffic pages.
- If you buy LLM-mention data, find out which scope your number is on before you quote it. Most people cannot answer this about their own dashboard.
- Record retrieval and citation separately. They are different events with different rates.
- Test a zero on the broadest scope, and look for a contrast case in the same session. A zero next to a 56 from the same call means something; a zero on its own might just be a filter.
- Store the parameters with the number — platform, locale, scope, date, subdomain handling.
Step 6 is the one we failed, and it cost us a published figure that was wrong by roughly four times. In five weeks you will not remember which scope you ran, and the number will be unusable. Five of these six steps are free.
Doing this without three APIs and a spreadsheet
Everything above came from a search console, a paid API called three times with different parameters, and a spreadsheet to join them — on one date, by hand. That is fine for an article. It is unworkable as a monitoring practice, which is the honest limitation of the method we have just handed you.
Semrush One is built for that continuity. Per its own product page, checked 2026-09-24, it offers to "See how your brand appears in Google, AI Overviews, and LLMs like ChatGPT", to "Monitor 317M+ relevant LLM prompts globally, including the largest US and ChatGPT databases", and to "Check your Site Health and AI Search Health scores side by side" — which is the cross-engine, cross-platform view our three instruments had to be stitched together to produce.
We have not used it. We are a Semrush affiliate, not a customer, and the description above is of Semrush's documentation rather than a review.
Semrush One
The four numbers in this article came from three instruments joined by hand on one date. Semrush One's stated purpose is that same cross-engine view, continuously.
See Semrush OneWhat we actually learned
The finding is not that we are cited thirteen times. Thirteen will be a different number by the time you read this, and so will the other three.
The finding is that four quantities the industry discusses as one moved independently on a single domain — and that we had spent five weeks quoting the wrong one of them under a confident label. A search console said one thing, a citation scope said another, a retrieval scope said a third, and Google's AI corpus said nothing at all.
Two opinions, offered as opinions. Where a first-party console and a scraped SERP disagree, believe the console. And a team that has never instrumented Bing is not measuring this dimension badly — they cannot see it at all.
Sources and method
| Figure | Source | Date |
|---|---|---|
Bing impressions, clicks, CTR, positions, /free, per-query | Bing Webmaster Tools, our own verified property | window 2026-04-10 → 09-18, pulled 2026-09-24 |
13 citations (sources scope) | DataForSEO LLM Mentions API, chat_gpt, 2840/en, subdomains included | 2026-09-24 |
| 56 any-scope appearances; 44 search-result appearances; Reddit 19 | Same endpoint, any scope | 2026-09-24 |
| Zero Google AI Overview mentions | Same endpoint, platform: google, broadest scope | 2026-08-19, 2026-09-22, 2026-09-24 |
| 47 ranked Google keywords, 0 in top 10, average position 67.3 | Third-party ranking data for our own domain | 2026-09-24 |
| ~87% of ChatGPT citations matching Bing's top 20 | LLMVLab, via Semrush's trend brief — not verified by us | 2026-04-21 |
| GPT-5.6 trusted-domain shortlisting | Semrush's trend brief and its sources — not verified by us | 2026-08-31 |
| 548,534 retrieved pages, 15,000 prompts, ~85% never cited | AirOps, reported by Search Engine Land — not reproduced by us | as published |
| Semrush One capabilities | Semrush's own product page | 2026-09-24 |
Written by Sohail Akhtar. Every figure here is a measurement on a stated date and should be read as one.
Some links may be affiliate links. We may earn a small commission at no extra cost to you.
Related Articles
Explore more guides and reviews from our experts.