Most AI search tools now sell you a single number. An AI visibility score, one figure, tracked over time, often averaged across four or five engines.
We run that monitoring ourselves, so I went and checked what the number is actually made of. Between 17 April and 27 September 2026 our system put 20,378 probes across five engines: ChatGPT, Google's AI Overviews, Perplexity, Gemini and Claude. Over the four weeks to 27 September it also recorded 26,467 individual source citations, every domain each engine pulled from, question by question, day by day.
The engines do not agree. Not roughly, not at the margins. On the same question on the same day, ChatGPT and Google's AI Overviews shared 8% of their sources.
One score cannot describe that. It averages five different games into a figure that moves for reasons you cannot act on.
Two of the five engines never showed us a link
Start with the biggest problem, because it is the one most scores quietly paper over.
Across 8,026 probes of Claude and Gemini, our monitoring captured a source URL four times.
That is not a failure of those models. It is what they are. Asked a question without a search tool attached, they answer from what they learned in training. They name businesses from memory. There is no link, because nothing was retrieved.
So a hit on Claude or Gemini and a hit on Perplexity are not the same event, and they should never be added together:
- Grounded engines (ChatGPT with search, Google's AI Overviews, Perplexity) retrieve live pages and cite them. A hit means your page was fetched and used. There is a click to win.
- Ungrounded models (Claude and Gemini answering without retrieval) draw on training data. A hit means the model already knows your name. There is no click, only recall.
In our data the recall numbers are low: Claude named the business in 8.3% of 4,428 probes, Gemini in 3.7% of 3,598. Combined, 6.2%.
Those two figures answer a completely different question from the other three. Whether a model remembers you is a function of how widely you were written about before its training cut-off. You cannot fix it this quarter by adding an FAQ block. Whether a grounded engine cites you is a retrieval problem you can work on this week.
Averaging them produces a number that is part memory test and part retrieval test, and tells you which lever to pull on neither.
The three engines that do cite barely overlap
For the three grounded engines we recorded every domain cited, for 277 question intents across 28 days.
When two engines answered the same question on the same day, here is how much of their source list they held in common:
| Engine pair | Shared sources | Average shared domains |
|---|---|---|
| ChatGPT vs Google AI Overviews | 8.0% | 1.23 |
| ChatGPT vs Perplexity | 8.4% | 1.94 |
| Google AI Overviews vs Perplexity | 17.4% | 4.07 |
ChatGPT and Google's AI Overviews, the two surfaces most businesses care about most, answered the same question from source lists with roughly one domain in common. In 42.7% of those cases they had no source in common at all.
This is not our finding alone. Ahrefs looked at the top 50 most mentioned websites on ChatGPT, Perplexity and AI Overviews across roughly 76.7 million AI Overviews and about 1.9 million prompts, and found 86% of them were not shared across the three. Only seven websites appeared in all three lists. Different data, different method, same conclusion.
The number of seats is not the same either
The engines also differ in how many sources they use at all.
| Engine | Median sources per answer | Average |
|---|---|---|
| Perplexity | 17 | 18.6 |
| Google AI Overviews | 9 | 10.2 |
| ChatGPT | 4 | 5.0 |
ChatGPT gives out about four seats per answer. Perplexity gives out seventeen.
That changes what a citation is worth and how hard it is to win. Being named among seventeen sources on Perplexity is a far easier target than being one of four on ChatGPT, and it is worth correspondingly less attention from the person reading. Treating both as one point of AI visibility flatters the easy win and hides the hard one.
Each engine trusts a different kind of source
Where the engines pull from differs just as sharply. Grouping every cited domain by type:
| Source type | ChatGPT | Google AI Overviews | Perplexity |
|---|---|---|---|
| Business or publisher site | 67.4% | 73.9% | 84.3% |
| Government or academic | 18.5% | 6.9% | 4.1% |
| Social, forum or video | 0.8% | 10.4% | 2.7% |
| Directory or marketplace | 1.4% | 3.5% | 4.0% |
| Platform's own properties | 11.9% | 5.3% | 5.0% |
ChatGPT leaned on government and academic sources more than twice as often as the other two, and almost never used social or forum content: 0.8%, against Google's 10.4%.
That last figure lines up with something the industry saw from the outside. Promptwatch tracked Reddit holding about 3.8% of ChatGPT citations for roughly three weeks, then dropping to about 0.5% on 14 August 2026. Our source window opens on 31 August, after that drop, and we measured 0.8% social across the board. Two independent measurements of the same change.
The practical read: a Reddit and forum strategy that earns citations in Google's AI Overviews is close to worthless on ChatGPT, while a cited government or standards reference does disproportionate work there. Directories were thin everywhere, between 1.4% and 4.0%.
Being named is not being linked
One more gap that a single score hides. On the grounded engines, the answer named the business far more often than it linked to it.
| Engine | Answers naming the business | Share that carried a link |
|---|---|---|
| Perplexity | 1,032 | 46.1% |
| ChatGPT | 1,028 | 37.7% |
| Google AI Overviews | 993 | 33.7% |
On ChatGPT, roughly six in ten mentions arrived with no clickable source. The business was recommended, and the reader had to go and find it.
That distinction is the whole revenue question. A mention builds recognition. A link delivers a visitor who has already been shortlisted. Any score that counts both as one point is measuring two different commercial outcomes as though they were the same.
What this costs
Put a price on it. In Australia a family law firm pays about A$210 for each enquiry it buys from Google search, across three firms' Google Ads accounts over the twelve months to September 2026 (PixelRush). An allied health clinic pays a median of A$107 per new-patient enquiry from Google Ads (Clinic Mastery Marketing, July 2026).
When a business optimises for one engine's preferences and reports a rising average score, the enquiries still go to whoever the other engines cite. The score goes up. The enquiries do not. And because the number is an average, nothing in it tells you which engine you just lost.
That is the failure mode worth naming: a single figure that rises while the thing it is supposed to stand for falls.
What to measure instead
Five changes, in order of how much they will change your decisions.
- Split retrieval from recall. Report grounded engines (ChatGPT, AI Overviews, Perplexity) separately from ungrounded ones (Claude, Gemini). They respond to different work on different timescales.
- Count links and mentions separately. A mention is awareness. A link is a visitor. Keep them in different columns.
- Score each engine on its own. One line per engine. An average across engines that share 8% of their sources is not a measurement.
- Adjust the target to the seat count. Being one of four on ChatGPT is not the same achievement as one of seventeen on Perplexity.
- Count enquiries at the door. Ask every new customer how they found you. It is still the only number a dashboard cannot flatter.
Anyone quoting you a single AI visibility figure should be able to show you the per-engine split underneath it. If they cannot, the number is decoration.
How we measured this, and what it does not show
The probes run daily against each engine's API, or via SERP retrieval for AI Overviews, checking a fixed set of buyer-intent questions for each client and recording whether the business was named, whether a source URL appeared, and every domain cited.
Sample: 20,378 probes across five engines, 17 April to 27 September 2026, and 26,467 source records across three grounded engines, 31 August to 27 September 2026. Twelve client accounts, ten industries, ten Australian businesses and two in the United States, mostly local service and franchise businesses.
The limits matter, so here they are plainly.
This is our own client book, not a sample of the web. These businesses are actively being optimised by us. Our citation rates of around 25% on grounded engines are not a market baseline and should not be read as one.
The cross-engine findings are the sturdier half. Selection bias inflates how often our clients get cited. It does not explain why two engines answering the same question pick different sources, use different numbers of them, or favour different source types. Those describe engine behaviour, and they are what I would take from this.
Claude and Gemini were probed without retrieval. Their consumer products can browse and cite. Our figures describe the models answering from training data, which is why the correct label is recall, not citation.
Engine answers are unstable between runs. Ronald Sielinski's analysis of repeated queries across three generative search platforms found citation rankings vary enough that single-run visibility metrics give a misleadingly precise picture (arXiv, final version 26 August 2026). Our numbers are aggregates over months for that reason, and any single day's reading is noisier than the totals suggest.
Measurement is getting harder. In September Google rolled out google.com/goto passthrough parameters on result URLs, explicitly to stop third-party tools reading its results. Anyone measuring AI Overviews from the outside, us included, should expect that to get less reliable, not more.
We are also running controlled format experiments to test which page structures actually win citations. The baselines are not complete, so there are no numbers to publish yet. When there are, they will be here with their limits attached.
The short version
There is no such thing as AI visibility. There are at least five surfaces with different retrieval behaviour, different numbers of source slots, different taste in sources, and two of them that do not hand out links at all.
Manage them as one number and you will optimise for the average of five games and win none of them.
Brain Buddy AI monitors all five engines separately for businesses in Australia and the United States, and works with partner agencies offering it to their own clients. To see your per-engine position, run the free AI visibility audit. Our client results, with sources, are on the results page.
Industries where an AI recommendation decides who gets the first enquiry have the most at stake: legal and accounting firms, healthcare and specialist practices, and trades and home services.
Sources
- Brain Buddy AI monitoring data, 17 April to 27 September 2026 (20,378 engine probes) and 31 August to 27 September 2026 (26,467 source citation records). Twelve client accounts across ten industries.
- Ahrefs, 86% of Top Mentioned Sources Are Not Shared Across ChatGPT, Perplexity, and AI Overviews, Patrick Stox, 12 June 2025 (~76.7M AI Overviews, 957k ChatGPT prompts, 953.5k Perplexity prompts, Ahrefs Brand Radar).
- Promptwatch data on Reddit's share of ChatGPT citations falling from about 3.8% to about 0.5% on 14 August 2026, as reported in ROI Revolution, September 2026 SEO & GEO News Recap.
- Ronald Sielinski, Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement, arXiv, final version 26 August 2026.
- Search Engine Roundtable, September 2026 Google Webmaster Report, on the
google.com/gotopassthrough parameters. - PixelRush, Law Firm Cost per Lead in Australia: 2026 Benchmarks.
- Clinic Mastery Marketing, Clinic Google Ads Cost Benchmarks 2026.
