Why one AI answer proves nothing

Ask an assistant the same question 100 times and there is under a 1 in 100 chance of getting the same list of brands back. Here is what that does to AI visibility measurement, and which numbers survive it.

Tom HardimanUpdated 17 August 20265 min read

Someone screenshots ChatGPT recommending a competitor and treats it as a finding. Someone else screenshots it recommending them and puts it on a slide. Both screenshots are close to worthless, and the reason is measurable.

The core finding

Grade A — provider-documented or controlled study

SparkToro ran 2,961 prompt runs with 600 volunteers across ChatGPT, Claude and Google's AI answers. There is less than a 1 in 100 chance that an identical prompt run 100 times returns the same list of brands, and less than a 1 in 1,000 chance of the same brands in the same order — even in tightly constrained categories.

Response length varied too, from two or three items to more than ten, with no discernible pattern. The assistants are not returning a ranking that occasionally wobbles. They are sampling from a distribution every single time.

Any tool that gives a 'ranking position in AI' is full of baloney.
Rand Fishkin, SparkToro, January 2026

Where the noise actually comes from

A 2026 variance-components study decomposed the non-determinism in brand answers. The result is uncomfortable for everyone selling in this category, including us.

Source of varianceShare
Re-running the same prompt34.8%
How the question is worded32.0%
Brand-in-context interaction29.6%
Which model you asked1.7%
Brand identity1.6%

Brand identity — the thing a client pays an agency to change — accounts for about 1.6% of the variance in what comes back. The reliability of a single answer as a measure of a brand's standing is approximately 0.01. That is not a rounding error, it is the whole problem.

The same study produced a finding that changes how sampling should be designed: adding diversity of wording and model reduces relative-error variance roughly fifteen times more than five additional repeats of the same prompt. Asking a different question is worth far more than asking the same one again.

It also sets a ceiling worth knowing. A design spanning eight languages, three models and fifteen paraphrases reaches a reliability of about 0.36. Better, and still modest. Nobody in this field is measuring precisely.

What this means for the tools

Fishkin's operational guidance is to run each prompt 60 to 100 times or more, implying roughly 20 runs across 7 phrasings and 7 models for accuracy within a percentage point. Set that against commercial tools offering 25-prompt or 15-prompt tiers.

Most AI visibility products are sampling one to two orders of magnitude below what their own headline precision implies. That does not make them useless. It makes their decimal places fictional.

Three traps that survive better sampling

Whoever picks the prompts picks the result

The same brand, on the same underlying data, can be shown at 20%, 16.8% or 31.4% share of voice depending on which prompts were chosen and which scoring formula was applied. A provider that selects its own prompt set and then reports improvement against that set has a plain conflict of interest. Ask to see the prompts, and ask whether they changed between reports.

The API is not the product

Answers pulled through an API reflect one account, one geography, one subscription tier and one moment of one model's state. Real users see a distribution shaped by personalisation, location, account state and constant product changes. Anyone measuring this way, ourselves included, is using a proxy and should say so.

Model drift looks exactly like your work

Providers ship model and retrieval updates constantly. One documented case saw GPT-4's accuracy on a fixed task fall from 84% to 51% in three months. A client's visibility can move materially in either direction for reasons entirely unrelated to anything an agency did — and it is genuinely difficult to tell the two apart.

The counter-argument, and what it does not answer

Agencies in this category have a standard response to the reproducibility problem, and it is worth engaging with properly because part of it is correct.

The argument runs: the criticism of tracked prompts is that they are invented questions rather than real conversations, and that objection is weakening. Prompt sets are increasingly built from Google Search Console queries, GA4 landing pages and on-site search terms, and some vendors now license anonymised real prompts from opt-in consumer panels. Scale is not a constraint either — practitioners track hundreds of prompts, not fifteen.

All of that is true, and a prompt set grounded in real buyer language is genuinely better than one somebody made up. But it answers a different objection to the one the research raises.

Better prompt sourcing fixes whether you are asking the right question. It does nothing about the fact that the same question, asked again, returns a different answer. Those are two separate problems, and only the first one has a purchasing solution.

The variance decomposition is what makes this stubborn. Re-running an identical prompt accounts for 34.8% of the variance on its own. Wording accounts for another 32%. Neither of those is addressed by sourcing the prompt more carefully — a perfectly sourced prompt still lands in a distribution.

Scale helps, and less than people assume. The same study found that a design spanning eight languages, three models and fifteen paraphrases reaches a reliability of about 0.36. That is a large, expensive panel, and it is still a long way from the precision implied by a dashboard reporting share of voice to one decimal place.

So the honest position sits between the two. Tracked prompts are not guesswork, and the people saying so are right. They are also not precise, and a report that gives you a number without a sample size and an acknowledgement of variance is overstating what it knows — whatever the prompts were built from.

So which numbers survive?

The line between a defensible measurement and theatre is sharper than the category admits.

DefensibleTheatre
Visibility as a share of runs mentioning you, across many prompts, each run repeatedly, with the sample size statedAny rank or position
Comparative visibility against a fixed competitor set on the same protocolShare of voice quoted to one decimal place
A trend over months, with model drift acknowledged as a confoundWeek-on-week movement, 'up two spots'
Presence or absence in the consideration setAttributing a lift to one specific blog post
'This change was undetectable at our sample size'A single screenshot as evidence of anything
Rankings? Probably nonsense. Visibility over many prompts and many runs? Surprisingly useful.
Rand Fishkin, SparkToro

How this shapes what we do

Two things follow directly, and both are visible in how our own reports are built.

  1. We spread runs across different wordings and different locations rather than repeating one question, because the variance research says diversity buys far more reliability than repetition does.
  2. We report counts out of a stated denominator and never a score out of 100. A count you can go and check yourself is more use than a composite number we invented, and given a single-answer reliability of 0.01, a score would be a decimal place laid over a coin flip.

It also sets the honest limit on what any report of this kind is. It is a record of what we observed on the dates stated, not a forecast, and running it again later will produce different figures. That is a property of the systems, not a defect in the method.

Common questions

Can you get ranked number one in ChatGPT?
No. There is no stable ranking to occupy. Asking an identical question 100 times returns the same list of brands less than 1% of the time, and the same brands in the same order less than 0.1% of the time. Any product reporting an AI rank is reporting noise.
How many times should a prompt be run to be meaningful?
Published guidance suggests 60 to 100 runs per prompt, and adding different wordings and different models buys far more reliability than repeating the same prompt. Most commercial tools sample one to two orders of magnitude below that.
Is AI visibility measurement worth anything at all?
Yes, within limits. Visibility measured as a share of runs across many prompts is useful and reasonably stable. Ranks, positions and share of voice quoted to one decimal place are not.

Sources

Every claim above should be checkable. Where a study has limits, they are stated rather than left out.

  1. SparkToro (Rand Fishkin) — 600 volunteers, 2,961 prompt runs, 12 prompts across ChatGPT, Claude and Google AI, January 2026Grade A. Large-sample reproducibility study.
  2. Žatuchin — 'Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers', arXiv:2607.13304, July 2026Grade A/B. Peer-review status varies; methodology published.
  3. Schulte, Bleeker & Kaufmann — 'Don't Measure Once: Measuring Visibility in AI Search (GEO)', arXiv:2604.07585, April 2026Grade A/B. Recommends bootstrap confidence intervals and repeated sampling.
  4. Chen, Zaharia & Zou — documented model drift on a fixed taskGrade B. Illustrates drift as a confound.

More on measuring ai visibility

  • Mentions and citations are not the same thing

    Being named in an answer and being used as a source are different events with different causes. In our own run, 77 companies were named and 204 domains were cited, and almost nobody managed both.

Want to know where you stand?

We ask five AI assistants for a company like yours and send you a free report showing how often you came up and who came up instead.

Get my free report