How it works

Three mechanisms drive variance. Sampling: unless temperature is zero, token selection is stochastic, and one different token early can change the whole list. Retrieval: when the system searches the live web, candidate sets shift with the index. Versioning and routing: providers update models and route requests across serving tiers, so behaviour changes without an announcement.

The practical implication is methodological. To say anything about visibility you need repetition — the same prompt many times, ideally across days and phrasings — and you have to report the variance, not just the modal answer.

Example

Ask an assistant for the leading tools in a crowded software category ten times and you may see a dozen distinct names, with only two or three appearing in most runs. Reporting only the first run would put a brand at the top of a category it appears in half the time. Our own repeated-prompt measurements are published as research notes with the raw data attached, precisely so this variance is visible rather than asserted.

Why it matters

Instability is the reason most "we rank #1 in ChatGPT" claims are unfalsifiable, and it is the reason serious measurement of AI visibility looks like polling: fixed prompts, repeated runs, reported confidence intervals, disclosed dates and model versions.