there’s a <1 in 100 chance that ChatGPT or Google’s AI, if asked 100X, will give you the same list of brands in any two responses
it’s more like 1 in 1,000 runs before you’d see two lists in the same order
any tool that gives a “ranking position in AI” is full of baloneyAnd a piece of advice to buyers that we are going to take literally:
stop throwing money at AI tracking products that don’t provide stats-backed, publicly-reviewable research. Before you spend a dime tracking AI visibility, make sure your provider answers the questions we’ve surfaced here and shows their math.This page is rekkal showing its math.
What the research means for a tracking tool
If any single answer is close to a coin flip, then:- One sample is not a measurement. A number from one run of one prompt tells you what the engine said once. It does not tell you what the engine says.
- A day-over-day chart is a noise chart. Most of the movement in a daily line is resampling, not the world changing.
- A rank position is not a fact. Order is less stable than membership, so a “position in AI” figure is the least trustworthy thing such a tool can print.
I now believe visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric
What rekkal does about it
1
Repeat the same prompts on a fixed schedule
Every prompt is re-run on every covered engine, every day. Your sample grows every
day whether or not anything changed, and the comparison is always like-for-like:
the same prompts, the same engines, the same slot length.The unit rekkal records — and bills — is one prompt × one engine × one daily slot. That
is not an implementation detail we happen to mention; it is the number every figure on
your dashboard is divided by, and it is printed on screen.
2
Count answers, not name-drops
An answer that mentions your brand five times is one observation. Counting each
mention separately would inflate the sample and make every confidence interval narrower
than the evidence justifies — which is how a tool ends up confidently reporting a
movement that is not there.
3
Put an exact confidence interval on every estimate
Clopper–Pearson exact binomial intervals, on the headline cards, on every per-engine bar,
on every row of the competitor ranking, and as a band on the charts. Small sample, wide
band — visibly.
4
Only claim a change when the intervals are disjoint
A direction and a colour are shown only when this period’s interval and the previous
period’s interval do not overlap. Everything else reads No measurable change, with the
point movement stated so you can see what was suppressed and why.This rule lives in one function, and the components that draw the arrows take their
colour from its verdict. A component cannot draw an arrow on noise even by mistake.
5
Give average position no direction at all
Because the research is clearest here, and because the arithmetic agrees: rekkal stores
no variance for a mean, so there is no interval to test with. Both periods are shown
side by side and the verdict is Change not testable. rekkal will not print a “your
rank improved” claim, on any plan, at any sample size.
6
Render absence as absence
A week that did not run has no point on the chart and nothing is interpolated across it.
A missing comparison names which half is missing. There is no zero-filling anywhere in
the read path — with one deliberate exception, below.
The one place rekkal writes a zero
Honesty about a rule includes its exception. If a week ran, and a brand in your tracked set was named in none of that week’s answers, that brand’s figure for that week is 0% — not “missing”. It was measured against a real denominator and went unmentioned, which is an observation. Everything else absent stays absent. A week with no runs contributes nothing at all: no zero, no point, no interpolation.What rekkal is measuring, precisely
- Provider APIs with grounding or web search enabled — OpenAI with web search, Perplexity Sonar, and Gemini with Google Search grounding. Which of these your plan covers is on Plans and limits.
- Not the consumer chat apps. No browser automation, no scraping. An API answer can differ from what the same provider’s app shows you, and rekkal says so rather than implying parity.
- One market and language per workspace, derived from your domain at setup. rekkal does not currently run the same prompt across several countries or languages in parallel.
Known limits of this approach
We would rather you read these here than discover them later.Daily cadence buys sample, not certainty
Daily cadence buys sample, not certainty
A daily slot means a genuine change takes days rather than weeks to become statistically
visible. That is the honest cost of not reporting noise. If you need to know within 24
hours that something moved, no sampling-based tool can honestly give you that — including
this one.
Sample size is bounded by your prompt allowance
Sample size is bounded by your prompt allowance
Your
n is prompts × engines × days. On a smaller plan, intervals stay wide for longer,
and rekkal will keep telling you so rather than pretending otherwise.An API is a proxy for a consumer surface
An API is a proxy for a consumer surface
rekkal measures a stable, comparable surface. It is not a transcript of what any
individual user sees in the ChatGPT app, and no tool that reports numbers can be.
Mentioned is not yet distinguished from recommended
Mentioned is not yet distinguished from recommended
A mention is counted whether the answer recommends you or lists you in passing. The
distinction matters and buyers ask for it; it is not built yet, and rekkal does not
imply it has it.
Only the visibility figure opens onto its answers
Only the visibility figure opens onto its answers
rekkal stores every answer and the exact position of every mention in it, and the
visibility card links straight to the answers it was counted from. Share of voice and
average position do not have that link yet — they are counted over the same stored
answers, but there is no screen that walks you through those two one answer at a time.