Your AI Competitor List Never Stops Growing
We nearly published a finding: AI names about 91 companies as the average brand's competitors, and only 18 recur. Then we sampled the same brands harder and 91 became 282, still climbing. The count was describing how often we asked, not the market. Here is the test that settles it, and what it means for every share-of-voice number.
We were about to publish a finding: AI engines name about 91 different companies as the average brand's competitors, and only around 18 of them recur. It was computed correctly across 429 brands and 135,996 competitor mentions.
It was also close to meaningless, and the reason is worth more than the original finding was. The number was not describing AI. It was describing how many times we had asked. Sample the same brands harder and 91 becomes 134, then 201, then 282, with no sign of stopping.
Which is the actual finding: you cannot count your AI competitor set. The count is a property of your measurement, not of your market, and that has a direct consequence for every share-of-voice number built on top of it.
Where the Number Came From
Most brands in our corpus are tracked lightly. Of 429, more than 360 have fewer than fifty responses that named a competitor at all. When you average a roster size across a group like that, you get a figure dominated by brands nobody has sampled much.
Worse, the "only 18 recur" half was close to circular. A company cannot appear three times in a brand with fourteen responses. Setting a recurrence threshold and then reporting how few names clear it, across brands that were barely sampled, measures the sampling rather than the market. The first version of this post had that backwards.
The Test That Settles It
Comparing lightly-tracked brands against heavily-tracked ones does not resolve this, because the two groups differ in more than sampling. A brand with hundreds of responses may simply sell into a broader market with more plausible rivals in it.
So we held the brand fixed instead. Take the fourteen brands with at least four hundred competitor-carrying responses, and recompute each one's roster using only its first ten responses, then its first twenty-five, and so on. Same brands, same market, same prompts. The only thing changing is how much of the data we let ourselves look at.
The Same Brands, Sampled Deeper
Fourteen brands, held constant, with the roster recomputed from their first N responses. At 50 responses the average brand has 91 named competitors. At 400 it has 282, and the line is still climbing.
Source: Ranqo production tracking data, within-brand rarefaction across the 14 brands with 400 or more competitor-carrying responses, measured July 2026. Depth N uses each brand's first N responses by record id, so the subsample is deterministic.
The roster grows from 33 companies to 282, and the curve has not flattened at the deepest point we can measure. The recurring set grows faster still, from 4 to 104, because names that looked like one-offs at shallow depth keep turning out to repeat.
Note where the original number lands. At fifty responses per brand, the average roster is exactly 91. That was never a fact about AI. It was a description of how deeply our cohort happened to be sampled, wearing the costume of a finding.
What Does Settle
Not everything runs away. Two quantities behave, and the difference between them and the raw count is what makes this usable.
Discovery Slows. It Does Not Stop.
Each extra response turns up fewer unseen companies, falling from 3.3 to 0.41. The share of names seen only once falls too, but appears to level off near 45% rather than heading for zero.
Source: same cohort and method. Left axis is the share of that brand's roster seen exactly once; right axis is new distinct names per additional response, computed between consecutive depths.
Discovery decelerates sharply. Early on, each additional response surfaces more than three companies you had not seen; by four hundred responses it surfaces about 0.4. You are still finding new names, just far more slowly, which is why the roster keeps climbing without ever exploding.
More interesting is the share of names seen exactly once. It falls steeply at first, from 71% to 54% by fifty responses, and then flattens, sitting at 45% at the deepest measurement. Some of that early drop is simply under-sampling correcting itself. What is left looks like a genuine floor: even when you look hard, close to half the companies an engine offers you appear once and never again.
What That Floor Actually Is
It would be convenient to call it hallucination. That is not honest, and the distinction matters if you are deciding what to ignore. The one-off group is at least four things mixed together, and they cannot be separated cleanly from name strings alone.
Under-sampling, still
Even at four hundred responses, some of those names would recur if you asked more. The curve is flattening, not finished. This is the explanation the first version of this post missed entirely.
Variants of companies already counted
A company written with its corporate suffix in one answer and without it in another looks like two entries. We canonicalize what we can and it does not catch everything.
Genuine but marginal rivals
Small or regional players do compete, and an engine may surface one for a narrow question and never again. Rare is not the same as wrong, though it does mean the company is not shaping how your market gets described.
Free association
Asked for a list, a model produces a list. Some names are plausible-sounding companies in roughly the right category that no strategist would accept as competitors.
How This Was Measured
| Field | Detail |
|---|---|
| Why rarefaction | Comparing thinly-tracked brands to heavily-tracked ones cannot separate sampling from the brand itself: a company with more tracking may simply operate in a broader market. Holding the brand fixed and varying only the depth removes that objection. |
| Cohort | The 14 brands with at least 400 responses that named a competitor. Every depth in the charts is these same 14 brands, so the lines describe sampling and nothing else. |
| Subsampling | For depth N we use each brand's first N qualifying responses, ordered by record id. Deterministic, so the figures reproduce exactly rather than shifting with each run. |
| A recurring name | A company named in three or more separate responses at that depth. Three is a reporting threshold, and the whole point of the chart is that what clears it depends on how far you sampled. |
| Known limits | Competitor names are free-text model output and canonicalization is imperfect, so some of the one-off group is variants of companies already counted. Fourteen brands is a small cohort, and they are Ranqo customers rather than a random sample. |
The honest limits: fourteen brands is a small cohort, and they are Ranqo customers rather than a random sample of the web. Four hundred responses is the deepest we can currently look, so we can say the curve has not flattened but not where it would. And canonicalization is imperfect, which inflates the one-off group by an amount we cannot quantify. The shape is the finding; treat the exact values as having slack. For how the corpus is built, our published study covers the methodology.
What to Do With This
- 01Stop treating the size of your AI competitor list as a metric. It measures your sampling, and it will rise every time you track more without telling you anything about the market.
- 02Hold your prompt set and cadence fixed, and record them next to any share-of-voice figure. Depth is part of the metric definition, not a footnote.
- 03When you do change the prompt set, restate the baseline rather than continuing the line. Otherwise the step change gets read as a market shift.
- 04Work from the recurring set, and re-derive it at your current depth rather than reusing a list built when you were sampling less often.
- 05Treat any published "AI names N competitors" claim, including the one we nearly made, as a statement about sample size until it says otherwise.
How Ranqo Helps
Ranqo runs a fixed prompt set on a fixed cadence, which is what makes a share-of-voice series comparable with itself over time, and records the full competitor set returned with every answer so the recurrence count is there rather than reconstructed. Share of voice is computed over the recurring set rather than the raw roster.
None of that makes the underlying number converge. It makes it consistent, which is the achievable goal. The companion finding from the same corpus is worth reading next: how often an engine recommends a brand with nothing in its source list to back it up.
Measure the same way every time
Ranqo tracks a fixed prompt set across every major engine, records who gets named alongside you, and keeps the depth constant so the trend means something. If you want the measurement method first, start with how to measure AI share of voice.
See your competitor setWritten by
Nisha Kumari
Nisha Kumari is Co-Founder at Ranqo, where she leads growth strategy and client acquisition. With a background in digital marketing and financial management, she specializes in SEO, Generative Engine Optimization, and helping brands build visibility across AI platforms.
Share this article
Related articles
How to Measure AI Share of Voice: The Three Decisions That Change the Number
The same prompt plays a different game on every platform -- Perplexity routinely stacks several times more citations into an answer than ChatGPT. Yet every guide hands you one formula: mentions divided by total, times 100. Share of voice is three decisions -- denominator, position weighting, aggregation -- before it's a number. This is the measurement playbook.
Why Every AI Visibility Tool Shows You a Different Number (Data From 102,025 Responses)
The same 102 brands scored 12% to 51.5% visibility depending only on which AI engine we asked (arXiv:2606.20065). Digiday's sources say three tools give three answers -- so are they all wrong? No: the disagreement is data. Tools differ because engines disagree, the target is probabilistic, and each defines 'visibility' differently. Here's the honest breakdown, and how to buy and read one anyway.
Recommendations Are Not Citations: What 19,588 AI Answers Show
Across 19,588 AI answers tracked for 430 brands, an engine recommended a brand 6,195 times and a cited page backed that up in only 1,963 of them. Being cited almost always means being recommended. Being recommended usually means no citation at all. Which makes a citation count a floor on your AI visibility rather than a measure of it.