Why one AI answer proves nothing
Ask an assistant the same question 100 times and there is under a 1 in 100 chance of getting the same list of brands back. Here is what that does to AI visibility measurement, and which numbers survive it.
Someone screenshots ChatGPT recommending a competitor and treats it as a finding. Someone else screenshots it recommending them and puts it on a slide. Both screenshots are close to worthless, and the reason is measurable.
The core finding
SparkToro ran 2,961 prompt runs with 600 volunteers across ChatGPT, Claude and Google's AI answers. There is less than a 1 in 100 chance that an identical prompt run 100 times returns the same list of brands, and less than a 1 in 1,000 chance of the same brands in the same order — even in tightly constrained categories.
Response length varied too, from two or three items to more than ten, with no discernible pattern. The assistants are not returning a ranking that occasionally wobbles. They are sampling from a distribution every single time.
“Any tool that gives a 'ranking position in AI' is full of baloney.”
Where the noise actually comes from
A 2026 variance-components study decomposed the non-determinism in brand answers. The result is uncomfortable for everyone selling in this category, including us.
| Source of variance | Share |
|---|---|
| Re-running the same prompt | 34.8% |
| How the question is worded | 32.0% |
| Brand-in-context interaction | 29.6% |
| Which model you asked | 1.7% |
| Brand identity | 1.6% |
Brand identity — the thing a client pays an agency to change — accounts for about 1.6% of the variance in what comes back. The reliability of a single answer as a measure of a brand's standing is approximately 0.01. That is not a rounding error, it is the whole problem.
The same study produced a finding that changes how sampling should be designed: adding diversity of wording and model reduces relative-error variance roughly fifteen times more than five additional repeats of the same prompt. Asking a different question is worth far more than asking the same one again.
It also sets a ceiling worth knowing. A design spanning eight languages, three models and fifteen paraphrases reaches a reliability of about 0.36. Better, and still modest. Nobody in this field is measuring precisely.
What this means for the tools
Fishkin's operational guidance is to run each prompt 60 to 100 times or more, implying roughly 20 runs across 7 phrasings and 7 models for accuracy within a percentage point. Set that against commercial tools offering 25-prompt or 15-prompt tiers.
Most AI visibility products are sampling one to two orders of magnitude below what their own headline precision implies. That does not make them useless. It makes their decimal places fictional.
Three traps that survive better sampling
Whoever picks the prompts picks the result
The same brand, on the same underlying data, can be shown at 20%, 16.8% or 31.4% share of voice depending on which prompts were chosen and which scoring formula was applied. A provider that selects its own prompt set and then reports improvement against that set has a plain conflict of interest. Ask to see the prompts, and ask whether they changed between reports.
The API is not the product
Answers pulled through an API reflect one account, one geography, one subscription tier and one moment of one model's state. Real users see a distribution shaped by personalisation, location, account state and constant product changes. Anyone measuring this way, ourselves included, is using a proxy and should say so.
Model drift looks exactly like your work
Providers ship model and retrieval updates constantly. One documented case saw GPT-4's accuracy on a fixed task fall from 84% to 51% in three months. A client's visibility can move materially in either direction for reasons entirely unrelated to anything an agency did — and it is genuinely difficult to tell the two apart.
The counter-argument, and what it does not answer
Agencies in this category have a standard response to the reproducibility problem, and it is worth engaging with properly because part of it is correct.
The argument runs: the criticism of tracked prompts is that they are invented questions rather than real conversations, and that objection is weakening. Prompt sets are increasingly built from Google Search Console queries, GA4 landing pages and on-site search terms, and some vendors now license anonymised real prompts from opt-in consumer panels. Scale is not a constraint either — practitioners track hundreds of prompts, not fifteen.
All of that is true, and a prompt set grounded in real buyer language is genuinely better than one somebody made up. But it answers a different objection to the one the research raises.
Better prompt sourcing fixes whether you are asking the right question. It does nothing about the fact that the same question, asked again, returns a different answer. Those are two separate problems, and only the first one has a purchasing solution.
The variance decomposition is what makes this stubborn. Re-running an identical prompt accounts for 34.8% of the variance on its own. Wording accounts for another 32%. Neither of those is addressed by sourcing the prompt more carefully — a perfectly sourced prompt still lands in a distribution.
Scale helps, and less than people assume. The same study found that a design spanning eight languages, three models and fifteen paraphrases reaches a reliability of about 0.36. That is a large, expensive panel, and it is still a long way from the precision implied by a dashboard reporting share of voice to one decimal place.
So the honest position sits between the two. Tracked prompts are not guesswork, and the people saying so are right. They are also not precise, and a report that gives you a number without a sample size and an acknowledgement of variance is overstating what it knows — whatever the prompts were built from.
So which numbers survive?
The line between a defensible measurement and theatre is sharper than the category admits.
| Defensible | Theatre |
|---|---|
| Visibility as a share of runs mentioning you, across many prompts, each run repeatedly, with the sample size stated | Any rank or position |
| Comparative visibility against a fixed competitor set on the same protocol | Share of voice quoted to one decimal place |
| A trend over months, with model drift acknowledged as a confound | Week-on-week movement, 'up two spots' |
| Presence or absence in the consideration set | Attributing a lift to one specific blog post |
| 'This change was undetectable at our sample size' | A single screenshot as evidence of anything |
“Rankings? Probably nonsense. Visibility over many prompts and many runs? Surprisingly useful.”
How this shapes what we do
Two things follow directly, and both are visible in how our own reports are built.
- We spread runs across different wordings and different locations rather than repeating one question, because the variance research says diversity buys far more reliability than repetition does.
- We report counts out of a stated denominator and never a score out of 100. A count you can go and check yourself is more use than a composite number we invented, and given a single-answer reliability of 0.01, a score would be a decimal place laid over a coin flip.
It also sets the honest limit on what any report of this kind is. It is a record of what we observed on the dates stated, not a forecast, and running it again later will produce different figures. That is a property of the systems, not a defect in the method.
Common questions
- Can you get ranked number one in ChatGPT?
- No. There is no stable ranking to occupy. Asking an identical question 100 times returns the same list of brands less than 1% of the time, and the same brands in the same order less than 0.1% of the time. Any product reporting an AI rank is reporting noise.
- How many times should a prompt be run to be meaningful?
- Published guidance suggests 60 to 100 runs per prompt, and adding different wordings and different models buys far more reliability than repeating the same prompt. Most commercial tools sample one to two orders of magnitude below that.
- Is AI visibility measurement worth anything at all?
- Yes, within limits. Visibility measured as a share of runs across many prompts is useful and reasonably stable. Ranks, positions and share of voice quoted to one decimal place are not.
Sources
Every claim above should be checkable. Where a study has limits, they are stated rather than left out.
- SparkToro (Rand Fishkin) — 600 volunteers, 2,961 prompt runs, 12 prompts across ChatGPT, Claude and Google AI, January 2026Grade A. Large-sample reproducibility study.
- Žatuchin — 'Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers', arXiv:2607.13304, July 2026Grade A/B. Peer-review status varies; methodology published.
- Schulte, Bleeker & Kaufmann — 'Don't Measure Once: Measuring Visibility in AI Search (GEO)', arXiv:2604.07585, April 2026Grade A/B. Recommends bootstrap confidence intervals and repeated sampling.
- Chen, Zaharia & Zou — documented model drift on a fixed taskGrade B. Illustrates drift as a confound.
More on measuring ai visibility
- Mentions and citations are not the same thing
Being named in an answer and being used as a source are different events with different causes. In our own run, 77 companies were named and 204 domains were cited, and almost nobody managed both.
Want to know where you stand?
We ask five AI assistants for a company like yours and send you a free report showing how often you came up and who came up instead.
Get my free report