Your AI Visibility Score Moves 15 Points for No Reason
We measured what an AI visibility score does when nothing happens. Between consecutive runs, with no fix shipped, the overall score moves up to 14.5 points and a single category 27.3. The narrower the slice, the noisier it gets, which is the opposite of what everyone assumes. Most reported wins in this category are sampling variance.
Run your AI visibility tracking twice, a week apart, and change nothing in between. The score will not be the same. Across 851 tracking runs we measured how far it moves on its own: the overall score swings up to 14.5 points, and a single category swings up to 27.3.
That is the number nobody in this category publishes, and it decides whether any reported improvement means anything. If your score went from 30 to 40 after a month of work, you have not learned that the work paid off. You have learned that your score is inside its own noise band.
This is not a claim about anyone's tool being sloppy, ours included. The movement comes from the engines and from sample size, so every platform measuring this way inherits it. The difference is only whether the tool tells you.
What We Measured
81 brands with at least 3 completed tracking runs each, 851 runs in total, and 27,803 individual results behind them. For every pair of consecutive runs we took the difference in the reported metric. No fix shipped in between, no prompt set edited, nothing done. Whatever moved, moved by itself.
One detail matters more than it looks. The metric is computed through the same function the product uses to show a customer their score, not through a cleaner formula written for the analysis. A hand-rolled version would disagree with what people actually see on screen, and then the noise floor would describe a number nobody is looking at.
The Noise Floor
How far a score moves when nothing changes
Consecutive completed runs across 81 brands and 851 runs, with no fix shipped in between. The bar is the 95th percentile of absolute change, in percentage points.
- Overall score
- 14.5 pts · SD 7.4
- One engine
- 20 pts · SD 10.9
- One category
- 27.3 pts · SD 13.1
The second figure is the standard deviation of the same comparisons. A narrower slice is noisier, not cleaner, because fewer prompts sit inside it.
Read the bars as: nineteen times out of twenty, a run-to-run change smaller than this is nothing. The overall score clears 14.5 points on its own often enough that a 10-point gain is not evidence of anything. The typical move is smaller, a standard deviation of 7.4 points, but the typical move is not the one you have to defend in a meeting.
We round these up into the thresholds the product enforces: 15 points overall, 20 for a single engine, 27 for a category. Below the line, a measured lift is not shown as a number at all. Showing one would turn a running total of sampling variance into something that looks like progress.
Narrower Is Noisier, Not Cleaner
The result that surprised us: the more specific the cut, the worse the noise. One category moves 1.9× as far as overall score does, 27.3 points against 14.5.
That runs against the instinct that a narrower slice is a sharper instrument. The mechanism is just sample size: a category holds a handful of prompts rather than the whole set, so a single answer flipping moves the average much further. The same applies to a single-engine view, which lands in between.
The practical consequence is uncomfortable. Category and per-engine breakdowns are the views people naturally drill into when they want to know what is working, and they are the views where a number is least likely to mean anything. It is one reason we publish a per-kind table rather than one global threshold.
The Individual Answers Are Fairly Stable
Here is the part that reframes the whole thing. Take a single prompt on a single engine and ask whether the verdict flipped between runs. It flipped 6.4% of the time across 24,467 comparable pairs. That is a fairly steady measurement.
So the answers are not chaotic. The score is. A visibility score is an average over a modest number of those cells, and averaging a small number of noisy-ish things produces something that jumps around more than any of its parts. The instability is arithmetic, not the engines being fickle.
It also tells you where to look. If the cells are stable and the score is not, then the useful question is never "did my number go up". It is which specific prompts changed hands, and whether they stayed changed.
The Two-Run Rule
Will a new mention still be there next run?
Share of (prompt, engine) cells that keep the mention into the following run, split by how many runs it has already survived.
- Seen once
- 54.3% · n=685
- Seen twice
- 96.1% · n=9,794
The mirror case: a cell absent for two runs running comes back on its own only 3.4% of the time (n=10,821).
A mention you have seen exactly once survives to the next run 54.3% of the time. A mention seen twice survives 96.1% of the time. That is a 41.8-point gap, and it is the single most useful thing in this post.
A first sighting is close to a coin flip. It is not a win, it is a candidate. The second consecutive sighting is what turns it into something you can report, plan around, or tell a client about. The rule that falls out is blunt: never count a new mention until you have seen it twice in a row.
The reverse holds just as strongly and is easier to act on. A prompt you have lost in two consecutive runs comes back on its own only 3.4% of the time. Those are not going to fix themselves, which makes them the cleanest work queue you have.
What a Ten-Point Gain Actually Means
Take a brand sitting at 30% visibility. A month later it reads 40%. That is the exact scenario an agency puts in a monthly report, and it is the scenario the measurement cannot support.
Ten points is below the 15-point band, which means a move that size happens on its own often enough that it carries no information. It might be real. The measurement simply cannot tell you which, and no amount of confidence in the underlying work changes that. Reported without a caveat, it is a claim the data does not make.
What would settle it is not a bigger number but another run. If the brand reads 40% again next cycle, and the specific prompts that gained are the same ones, the picture changes completely: two consecutive readings of the same thing is a different piece of evidence from one reading twice as large. This is the difference between measuring something and hoping.
There is an uncomfortable corollary for anyone reporting monthly on a weekly cadence. Four runs is not many samples, and a month is short enough that most genuine content work has not landed yet. The honest monthly report often says the number moved and we cannot yet attribute it, alongside the list of prompts that changed hands. That reads as weaker than a confident percentage, and it is worth considerably more, because the alternative is claiming a win that reverses next month.
The same logic cuts the other way and is easier to sell. A month where the score fell eight points is not a month where anything went wrong. Being able to say so, and to point at a threshold you set in advance rather than one chosen after seeing the result, is most of what separates a measurement from a narrative.
What to Do With This
- 01Put a threshold on your reporting before you look at the number. Ours is 15 points overall and 27 for a category. Anything smaller is "no detectable change", not a small win.
- 02Count runs, not weeks. Two consecutive runs is the minimum before a new mention is real, and no amount of calendar time substitutes for a second measurement.
- 03Track which prompts changed, not just the headline. The cells are stable enough to be worth reading individually, which the score is not.
- 04Prefer ratios to levels where you can. Share of voice moves less than a raw visibility level, because a run that returns more names inflates the denominator too.
- 05Be suspicious of a tool that reports a two-point improvement with a straight face. That is also part of why two tools disagree about the same brand on the same day.
We wrote the qualitative version of this argument months ago, in a post about why one measurement is never enough. It leaned on the published study and made the case without a number for how much movement to expect. This is that number.
See your own movement, run over run
Ranqo re-runs the same prompt set on a fixed cadence and holds a measured lift back until it clears the noise band, so a quiet week reads as a quiet week rather than as progress. The free checker runs unbranded prompts from your category against ChatGPT live. No signup.
Run the free checkWritten by
Nisha Kumari
Nisha Kumari is Co-Founder at Ranqo, where she leads growth strategy and client acquisition. With a background in digital marketing and financial management, she specializes in SEO, Generative Engine Optimization, and helping brands build visibility across AI platforms.
Share this article
Related articles
Why You Can't Measure AI Visibility Just Once
Run the same AI query twice and you may get a different answer. So is AI visibility just noise? Mostly not: in our 102-brand study, 77.5% of brand-prompt-engine cells were deterministic, always cited or never. But sentiment flips 6.7x more often than whether you're mentioned. The fix is boring and it works: measure the same prompts repeatedly.
Why Every AI Visibility Tool Shows You a Different Number (Data From 102,025 Responses)
The same 102 brands scored 12% to 51.5% visibility depending only on which AI engine we asked (arXiv:2606.20065). Digiday's sources say three tools give three answers -- so are they all wrong? No: the disagreement is data. Tools differ because engines disagree, the target is probabilistic, and each defines 'visibility' differently. Here's the honest breakdown, and how to buy and read one anyway.
How to Measure AI Share of Voice: The Three Decisions That Change the Number
The same prompt plays a different game on every platform -- Perplexity routinely stacks several times more citations into an answer than ChatGPT. Yet every guide hands you one formula: mentions divided by total, times 100. Share of voice is three decisions -- denominator, position weighting, aggregation -- before it's a number. This is the measurement playbook.