Skip to main content
Ranqo
Features
Search VisibilityMonitor how AI sees your brand — visibility, position, and sentimentPrompt IntelligenceThe exact questions buyers ask AI — and whether you're in the answerCompetitor BenchmarkingSee who AI recommends instead of you — share of voice, position, and gapsSource AnalyticsThe domains AI trusts when it answers — and how to become one
Visitor AnalyticsSee every AI bot and human visitor — the traffic GA4 missesAction CenterKnow what to do next — every action ranked by impact, tracked to donePage OptimizationSix-dimension audits that show why AI cites a page — or skips itContent LabFrom visibility gap to published post — content AI wants to cite
One platform. Every AI conversation, captured.Compare all
Pricing
Resources
BlogResearch & playbooks for AI searchResearchPublished papers on AI searchAI Visibility ReportYour AEO/GEO baseline, emailedFree ToolsAudit or generate, no signupCompare GEO/AEO platformsRanqo vs alternatives, head-to-head
Features
Search VisibilityMonitor how AI sees your brand — visibility, position, and sentimentPrompt IntelligenceThe exact questions buyers ask AI — and whether you're in the answerCompetitor BenchmarkingSee who AI recommends instead of you — share of voice, position, and gapsSource AnalyticsThe domains AI trusts when it answers — and how to become oneVisitor AnalyticsSee every AI bot and human visitor — the traffic GA4 missesAction CenterKnow what to do next — every action ranked by impact, tracked to donePage OptimizationSix-dimension audits that show why AI cites a page — or skips itContent LabFrom visibility gap to published post — content AI wants to cite
PricingResources
BlogResearch & playbooks for AI searchResearchPublished papers on AI searchAI Visibility ReportYour AEO/GEO baseline, emailedFree ToolsAudit or generate, no signupCompare GEO/AEO platformsRanqo vs alternatives, head-to-head
Book a demoLog inGet Started
Research

Your AI Visibility Score Moves 15 Points for No Reason

We measured what an AI visibility score does when nothing happens. Between consecutive runs, with no fix shipped, the overall score moves up to 14.5 points and a single category 27.3. The narrower the slice, the noisier it gets, which is the opposite of what everyone assumes. Most reported wins in this category are sampling variance.

Nisha Kumari·Aug 24, 2026·8 min read

On this page

Summarize with AI

ChatGPTChatGPTClaudeClaudePerplexityPerplexity

Run your AI visibility tracking twice, a week apart, and change nothing in between. The score will not be the same. Across 851 tracking runs we measured how far it moves on its own: the overall score swings up to 14.5 points, and a single category swings up to 27.3.

That is the number nobody in this category publishes, and it decides whether any reported improvement means anything. If your score went from 30 to 40 after a month of work, you have not learned that the work paid off. You have learned that your score is inside its own noise band.

This is not a claim about anyone's tool being sloppy, ours included. The movement comes from the engines and from sample size, so every platform measuring this way inherits it. The difference is only whether the tool tells you.

What We Measured

81 brands with at least 3 completed tracking runs each, 851 runs in total, and 27,803 individual results behind them. For every pair of consecutive runs we took the difference in the reported metric. No fix shipped in between, no prompt set edited, nothing done. Whatever moved, moved by itself.

One detail matters more than it looks. The metric is computed through the same function the product uses to show a customer their score, not through a cleaner formula written for the analysis. A hand-rolled version would disagree with what people actually see on screen, and then the noise floor would describe a number nobody is looking at.

The Noise Floor

How far a score moves when nothing changes

Consecutive completed runs across 81 brands and 851 runs, with no fix shipped in between. The bar is the 95th percentile of absolute change, in percentage points.

Overall score
14.5 pts · SD 7.4
One engine
20 pts · SD 10.9
One category
27.3 pts · SD 13.1

The second figure is the standard deviation of the same comparisons. A narrower slice is noisier, not cleaner, because fewer prompts sit inside it.

Read the bars as: nineteen times out of twenty, a run-to-run change smaller than this is nothing. The overall score clears 14.5 points on its own often enough that a 10-point gain is not evidence of anything. The typical move is smaller, a standard deviation of 7.4 points, but the typical move is not the one you have to defend in a meeting.

We round these up into the thresholds the product enforces: 15 points overall, 20 for a single engine, 27 for a category. Below the line, a measured lift is not shown as a number at all. Showing one would turn a running total of sampling variance into something that looks like progress.

Narrower Is Noisier, Not Cleaner

The result that surprised us: the more specific the cut, the worse the noise. One category moves 1.9× as far as overall score does, 27.3 points against 14.5.

That runs against the instinct that a narrower slice is a sharper instrument. The mechanism is just sample size: a category holds a handful of prompts rather than the whole set, so a single answer flipping moves the average much further. The same applies to a single-engine view, which lands in between.

The practical consequence is uncomfortable. Category and per-engine breakdowns are the views people naturally drill into when they want to know what is working, and they are the views where a number is least likely to mean anything. It is one reason we publish a per-kind table rather than one global threshold.

The Individual Answers Are Fairly Stable

Here is the part that reframes the whole thing. Take a single prompt on a single engine and ask whether the verdict flipped between runs. It flipped 6.4% of the time across 24,467 comparable pairs. That is a fairly steady measurement.

So the answers are not chaotic. The score is. A visibility score is an average over a modest number of those cells, and averaging a small number of noisy-ish things produces something that jumps around more than any of its parts. The instability is arithmetic, not the engines being fickle.

It also tells you where to look. If the cells are stable and the score is not, then the useful question is never "did my number go up". It is which specific prompts changed hands, and whether they stayed changed.

The Two-Run Rule

Will a new mention still be there next run?

Share of (prompt, engine) cells that keep the mention into the following run, split by how many runs it has already survived.

Seen once
54.3% · n=685
Seen twice
96.1% · n=9,794

The mirror case: a cell absent for two runs running comes back on its own only 3.4% of the time (n=10,821).

A mention you have seen exactly once survives to the next run 54.3% of the time. A mention seen twice survives 96.1% of the time. That is a 41.8-point gap, and it is the single most useful thing in this post.

A first sighting is close to a coin flip. It is not a win, it is a candidate. The second consecutive sighting is what turns it into something you can report, plan around, or tell a client about. The rule that falls out is blunt: never count a new mention until you have seen it twice in a row.

The reverse holds just as strongly and is easier to act on. A prompt you have lost in two consecutive runs comes back on its own only 3.4% of the time. Those are not going to fix themselves, which makes them the cleanest work queue you have.

What a Ten-Point Gain Actually Means

Take a brand sitting at 30% visibility. A month later it reads 40%. That is the exact scenario an agency puts in a monthly report, and it is the scenario the measurement cannot support.

Ten points is below the 15-point band, which means a move that size happens on its own often enough that it carries no information. It might be real. The measurement simply cannot tell you which, and no amount of confidence in the underlying work changes that. Reported without a caveat, it is a claim the data does not make.

What would settle it is not a bigger number but another run. If the brand reads 40% again next cycle, and the specific prompts that gained are the same ones, the picture changes completely: two consecutive readings of the same thing is a different piece of evidence from one reading twice as large. This is the difference between measuring something and hoping.

There is an uncomfortable corollary for anyone reporting monthly on a weekly cadence. Four runs is not many samples, and a month is short enough that most genuine content work has not landed yet. The honest monthly report often says the number moved and we cannot yet attribute it, alongside the list of prompts that changed hands. That reads as weaker than a confident percentage, and it is worth considerably more, because the alternative is claiming a win that reverses next month.

The same logic cuts the other way and is easier to sell. A month where the score fell eight points is not a month where anything went wrong. Being able to say so, and to point at a threshold you set in advance rather than one chosen after seeing the result, is most of what separates a measurement from a narrative.

What to Do With This

  • 01Put a threshold on your reporting before you look at the number. Ours is 15 points overall and 27 for a category. Anything smaller is "no detectable change", not a small win.
  • 02Count runs, not weeks. Two consecutive runs is the minimum before a new mention is real, and no amount of calendar time substitutes for a second measurement.
  • 03Track which prompts changed, not just the headline. The cells are stable enough to be worth reading individually, which the score is not.
  • 04Prefer ratios to levels where you can. Share of voice moves less than a raw visibility level, because a run that returns more names inflates the denominator too.
  • 05Be suspicious of a tool that reports a two-point improvement with a straight face. That is also part of why two tools disagree about the same brand on the same day.

We wrote the qualitative version of this argument months ago, in a post about why one measurement is never enough. It leaned on the published study and made the case without a number for how much movement to expect. This is that number.

See your own movement, run over run

Ranqo re-runs the same prompt set on a fixed cadence and holds a measured lift back until it clears the noise band, so a quiet week reads as a quiet week rather than as progress. The free checker runs unbranded prompts from your category against ChatGPT live. No signup.

Run the free check

Written by

Nisha Kumari

Co-Founder at Ranqo

Nisha Kumari is Co-Founder at Ranqo, where she leads growth strategy and client acquisition. With a background in digital marketing and financial management, she specializes in SEO, Generative Engine Optimization, and helping brands build visibility across AI platforms.

On this page

Summarize with AI

ChatGPTChatGPTClaudeClaudePerplexityPerplexity

Share this article

Related articles

GuideJun 21, 2026

Why You Can't Measure AI Visibility Just Once

Run the same AI query twice and you may get a different answer. So is AI visibility just noise? Mostly not: in our 102-brand study, 77.5% of brand-prompt-engine cells were deterministic, always cited or never. But sentiment flips 6.7x more often than whether you're mentioned. The fix is boring and it works: measure the same prompts repeatedly.

4 min readRead article
ResearchJul 7, 2026

Why Every AI Visibility Tool Shows You a Different Number (Data From 102,025 Responses)

The same 102 brands scored 12% to 51.5% visibility depending only on which AI engine we asked (arXiv:2606.20065). Digiday's sources say three tools give three answers -- so are they all wrong? No: the disagreement is data. Tools differ because engines disagree, the target is probabilistic, and each defines 'visibility' differently. Here's the honest breakdown, and how to buy and read one anyway.

11 min readRead article
GuideJun 11, 2026

How to Measure AI Share of Voice: The Three Decisions That Change the Number

The same prompt plays a different game on every platform -- Perplexity routinely stacks several times more citations into an answer than ChatGPT. Yet every guide hands you one formula: mentions divided by total, times 100. Share of voice is three decisions -- denominator, position weighting, aggregation -- before it's a number. This is the measurement playbook.

9 min readRead article
Ranqo

Ranqo is the complete AI search visibility suite: track, analyze, and improve how your brand shows up across ChatGPT, Claude, Perplexity, Gemini & Grok.

Read our reviews on G2Ranqo approved on SaasHub

Product

  • Search Visibility
  • Prompt Intelligence
  • Competitor Benchmarking
  • Source Analytics
  • Visitor Analytics
  • Action Center
  • Page Optimization
  • Content Lab

Company

  • About
  • Pricing
  • Book a demo
  • Contact

Legal

  • Privacy
  • Terms
  • Cookies

Resources

  • Blog
  • Research
  • Compare
  • AI Visibility Report

Free Tools

  • All Free Tools
  • AI Visibility Checker
  • AI Readiness Score
  • LLMs.txt Generator
  • Robots.txt Generator
  • Page Token Inspector
  • AI Content Grader
  • AI Search Crawler Inspector
  • AI Query Fan-Out Generator

© 2026 Ranqo. All rights reserved.