← Unmentioned

Why do AI answers change every time you ask the same question?

Language models sample from a probability distribution rather than looking up a fixed answer, and search-grounded engines fetch different sources each time — so identical prompts return different brand lists. Any AI visibility figure based on a single run is therefore a sample, not a measurement.

Two sources of variation, stacked

The model itself is probabilistic: at each step it samples from a distribution over possible next words, so two runs of the same prompt genuinely diverge rather than repeating.

On top of that, engines that search the web fetch a live result set that changes between runs. Different sources in, different brands out. The two effects compound.

Why this breaks 'average position'

Most AI visibility tools ask each prompt once per day and report an average position across prompts. That number has a precision it has not earned: an average over single samples of a high-variance process is not a measurement of anything stable.

It also produces the failure mode customers actually feel — the score moves, you cannot tell whether anything changed, and the tool cannot tell you either.

The honest response is not to hide the variance but to measure it. Ask each question several times, report how often the brand appeared rather than where it ranked once, and label results that came from a single run as what they are.

What to do about it

Treat repetition as the unit of measurement. A brand named in 3 of 3 runs and a brand named in 1 of 3 runs are in genuinely different positions, even though a single-sample tool would report both as 'mentioned'.

And be suspicious of any tool that will not tell you how many times it asked. If the methodology isn't stated, the number isn't checkable.

Check your own domain

Free, no signup, about two minutes.

Free · No signup · ~2 minutes

Common questions

Does setting temperature to zero fix it?

Not for this purpose. It reduces sampling variation but does not remove it in practice, and it does nothing about the larger source — the changing set of web results a grounded engine retrieves. It also stops reflecting what real users see, since consumer assistants do not run at zero.

How many runs are enough?

Three is enough to tell a stable result from a volatile one, which is the distinction that changes what you should do. More runs tighten the rate but cost proportionally more.

So is AI visibility tracking meaningless?

No — it means the unit has to be a rate over repeated runs rather than a position from one. Tracked over time with a consistent method, that rate is a real signal. A single number with no stated methodology is not.

Related questions

Last updated 24 August 2026