Why the five AI assistants disagree about who to recommend
We measured every source five assistants used to answer the same questions. Not one domain was cited by all five, and 89% were cited by only one. Our own original data on why a single AI visibility number is misleading.
Most AI visibility reporting produces a single number. One score, one share of voice, one position. That framing assumes the assistants are broadly looking at the same web and broadly agreeing about it. We measured whether that is true, using our own audit tool, and the answer is not close.
What we did
We ran our own business through our own audit: three questions a buyer in our category actually asks, two phrasings each, across ChatGPT, Claude, Gemini, Google AI Overviews and Perplexity. Twenty-nine usable answers. We then recorded every source domain each assistant cited.
The point of auditing ourselves is that we can publish the whole thing. We are not describing a client, and nothing here is anonymised.
Not one source was common to all five
Across 29 answers the five assistants cited 204 distinct domains between them. Zero of those domains were cited by all five assistants. 182 of the 204 — 89% — were cited by exactly one assistant and no other.
The pairwise overlaps are as low as the headline suggests. Jaccard similarity measures shared domains as a share of the combined set; 1.0 would be identical source lists and 0 would be no shared sources at all.
| Pair | Shared domains | Similarity |
|---|---|---|
| Google AI Overviews × Perplexity | 6 | 0.072 |
| Claude × Google AI Overviews | 7 | 0.071 |
| Claude × Perplexity | 6 | 0.055 |
| Gemini × Perplexity | 5 | 0.050 |
| Claude × Gemini | 5 | 0.043 |
| Gemini × Google AI Overviews | 4 | 0.043 |
| ChatGPT × Claude | 2 | 0.023 |
| ChatGPT × Gemini | 1 | 0.013 |
| ChatGPT × Google AI Overviews | 0 | 0.000 |
| ChatGPT × Perplexity | 0 | 0.000 |
ChatGPT and Perplexity did not share a single source. Neither did ChatGPT and Google's AI Overviews. The best-agreeing pair in the whole set overlapped on 7% of their combined sources.
They do not even read the same amount
| Assistant | Distinct domains cited across 29 answers |
|---|---|
| Claude | 66 |
| Gemini | 56 |
| Perplexity | 49 |
| Google AI Overviews | 40 |
| ChatGPT | 22 |
Claude drew on three times as many distinct sources as ChatGPT for the same questions. If you were judging your visibility from ChatGPT alone you would be looking at the narrowest source pool of the five, and concluding something about the web from a keyhole.
Why this happens
The assistants are not variations on one search engine. They run on different indexes, and that is the mechanical reason the source lists barely intersect.
- ChatGPT discovers URLs through the Bing Search API, so Microsoft's index decides what is even eligible.
- Google AI Overviews, AI Mode and Gemini are grounded in Google's own Search ranking and quality systems.
- Perplexity runs its own independent crawl and index.
- Microsoft Copilot is Bing-native, sharing ChatGPT's discovery gate.
Even within one company the surfaces disagree. Ahrefs compared 540,000 query pairs and found Google's AI Mode and AI Overviews cited the same URLs only 13.7% of the time, while reaching broadly the same conclusions about 86% of the time. One company, one index, two different answers about whose page to trust.
Being retrieved is not being cited
There is a second filter after discovery, and it is severe. AirOps analysed 548,534 pages across 15,000 prompts and found ChatGPT cites roughly 15% of the pages it retrieves. The other 85% are pulled in, assessed and discarded.
That reframes the work. Getting found is the easy half. Being kept is the half that decides whether you are named, and it depends on what is on the page once the assistant has it open.
What this means for measurement
- A single AI visibility score averages five populations that share almost no sources. The average describes none of them.
- Per-assistant work barely transfers. If your source overlap between two engines is under 10%, fixing your presence on the sources one engine reads does very little for the other.
- Judge coverage by which assistants you were measured on. A tool checking one or two engines is not measuring your AI visibility, it is measuring one index.
- Expect the picture to differ most where it matters most. ChatGPT, the assistant with the largest audience, drew on the fewest sources in our run.
The limits of this measurement
This is 29 answers, in one category, on one day, from one set of accounts and one geography. It is enough to show that the source lists barely intersect — a result that stark does not come from sampling noise — but it is not a precise estimate of the overlap, and we would not present it as one. Anyone quoting a figure like this to a decimal place is overstating what a run of this size can support.
It is also a single run. AI answers are non-deterministic, so repeating this would produce different domains and somewhat different overlaps. The direction is the finding; the exact numbers are one observation.
Common questions
- Do all AI assistants use the same sources?
- No, and the gap is much larger than most reporting implies. In our own measured run across 29 answers, the five major assistants cited 204 distinct domains and not one domain was cited by all five. 89% were cited by only one assistant. ChatGPT and Perplexity shared no sources at all.
- Why does ChatGPT recommend different businesses to Gemini?
- They run on different indexes. ChatGPT discovers URLs through Bing, Gemini and AI Overviews are grounded in Google Search, and Perplexity runs its own crawl. Different index, different candidate pages, different answer.
- Is a single AI visibility score useful?
- Not on its own. It averages engines whose source lists barely intersect, so the average describes none of them accurately. Per-assistant counts against a stated denominator are more honest and more actionable.
Sources
Every claim above should be checkable. Where a study has limits, they are stated rather than left out.
- Sentinay — original measurement, 29 answers across five assistants, 17 August 2026. Full raw data retained, including every cited URL.Grade B. Our own single run; small sample, stated limits in the guide.
- Ahrefs — AI Mode versus AI Overviews citation comparison, 540,000 query pairsGrade B. Large-sample observational.
- AirOps — 548,534 pages across 15,000 prompts, retrieval versus citation rateGrade B. Vendor-published, large sample.
- Perplexity Help Center — independent index and crawler documentationGrade A. Provider documentation.
- Google Search Central — AI features grounded in core ranking systemsGrade A. Provider documentation.
More on getting cited in ai answers
- The source ecosystem: how AI actually decides who to name
When we asked five assistants who to hire in our own category, 59 of 108 citations were other people's roundup lists. Five of the twelve most-named companies sat on one directory. Being named is mostly about other people's pages.
- What actually works in AI visibility, graded by evidence
Twenty interventions sold as AI visibility work, each rated by how strong the evidence behind it actually is. Three are gates you must pass. One has a rigorous causal study behind it, and that study found nothing.
- How AI assistants pick which businesses to name
ChatGPT, Perplexity, Google's AI answers and Copilot use four different pipelines to decide who gets recommended. An intervention that works on one can do nothing on another, which is why they have to be measured separately.
- What to fix first for AI search
A running order based on what the evidence supports rather than what is easiest to sell. Three gates before anything else, then off-site work, then your own pages — which come later than almost every agency will tell you.
- Does llms.txt do anything?
Over 844,000 sites have adopted llms.txt. No major AI provider has committed to reading it in production retrieval, and the largest correlation study found no relationship with citations. Here is what it is actually for.
Want to know where you stand?
We ask five AI assistants for a company like yours and send you a free report showing how often you came up and who came up instead.
Get my free report