블로그로 돌아가기

Measurement

A confidence interval on one AI run is fiction

The same question, same engine, same day can produce different answers. That is the nature of AI search, not a bug. The 3 rules we use to avoid showing you fake statistics, and how to read a visibility score without fooling yourself.

작성자: WhoCanFindMe5분 소요
Notebook chart showing five scattered readings from repeated runs of one AI question, with a single tight error bar crossed out.

Are AI visibility tools accurate? It is the right question to ask, and the honest answer from a company that builds one is: not in the way most dashboards imply. AI answers change between runs. The same question, put to the same engine on the same day, can name your business once and skip it the next time. A tool that hides that wobble behind a tight confidence interval is not measuring more carefully than its rivals. It is inventing precision. This post explains where the wobble comes from, the 3 rules we bind ourselves to so we never fake it, and how to read a visibility score without fooling yourself.

Are AI visibility tools accurate? The scepticism is deserved

Marketers have started saying in public what many suspected in private. Digiday reported in May 2026 that inconsistent results are fuelling scepticism about expensive AI visibility tools, with spending on them described as a "necessary evil". The pattern behind the frustration is simple. A brand runs the same check twice, gets 2 different numbers, and reasonably asks which one was true.

The scepticism is deserved. But the usual conclusion (the tools must be broken) is slightly off target. The measurements vary because the thing being measured varies. The variance is not the failure. The failure is any tool that pretends the variance is not there.

Why the same question gets different answers

An AI engine does not look up a stored answer the way a database does. It generates a fresh answer every time, and a small amount of randomness is built into that process on purpose. Statisticians call this behaviour stochastic: the same input can produce different outputs, by design. All 4 engines we test (ChatGPT, Perplexity, Gemini, Claude) work this way.

Randomness is not the only source of movement. Engines that search the live web can retrieve different pages on different runs. Models are updated without notice. A small change in how a buyer phrases a question changes which sources get pulled into the answer. None of this is a defect you can report and wait for a fix. It is the nature of the instrument.

The practical consequence: measuring AI visibility is closer to measuring wind than measuring a wall. One gust tells you very little. A month of readings tells you the prevailing direction.

The 3 rules that keep fiction out of our reports

A confidence interval is the range a statistician draws around a number to show how much it would wobble if you measured again. Drawing one honestly requires repeated measurements. Drawing one from a single run is not statistics. It is decoration.

So we follow 3 suppression rules. A suppression rule means we hide a statistic entirely rather than show a version of it we cannot defend. All 3 are documented on our methodology page.

1. No stability bands below 3 repetitions

A stability band shows how much an answer moves when we ask the same question again. With fewer than 3 repetitions we cannot know that, so we show no band at all. A blank is more honest than a guess.

2. No representation accuracy below 5 samples

Representation accuracy is our measure of whether the engines describe your business correctly when they do mention it. Below 5 samples that percentage is a coin flip dressed up as a finding, so we refuse to display it.

3. A single-run score is a reading, not a truth

We label a first scan for what it is: one reading from a moving system. It is genuinely useful, the way a first blood pressure reading is useful. No doctor would diagnose you from it alone, and nobody should rebuild a website over one scan.

If a tool shows you a tight confidence interval from one run, it is showing you fiction. Ask the vendor how many repetitions sit behind the band. The answer tells you most of what you need to know about the rest of the product.

How to read a visibility score responsibly

Four habits make a single score useful instead of misleading.

Treat it as a snapshot. It tells you roughly where you stand today, not where you will stand tomorrow.

Never react to one run. A bad reading is not a crisis and a good one is not a victory. Change something only when repeated readings agree.

Compare like with like. A trend only means something if the question set, the engines and the site being scanned stay the same between readings. Our own audit taught us this the hard way: scanning a local test copy of a site instead of the live one wrecks the comparison on its own.

Check what got counted. A raw mention count flatters everyone. What matters is whether the engine actually talked about your business, and whether it recommended you or merely listed you. We wrote about those distinctions in mentions versus recommendations and the three kinds of AI mention.

When AI rank tracking becomes reliable

Is AI rank tracking reliable? On a single run, no, and no vendor can make it so. It becomes reliable the way any measurement of a noisy system does: through repetition over time.

A real trend looks like this. The same question set, run weekly, for at least a month. The weekly numbers wander, because they always will. Reliability is the direction holding across weeks and across engines. If your visibility is genuinely improving, the improvement shows up in ChatGPT and Perplexity and Gemini and Claude across several readings, not in one lucky Tuesday.

And the honest caveat: even a trend carries noise. We do not promise rankings and we do not guarantee citations, because nobody controls what these engines say. What weekly tracking gives you is the earliest trustworthy signal that the work you are doing (fixing crawler access, being named on pages the engines read, keeping your details consistent) is moving the needle.

Start with a reading, act on the trend

A single scan cannot tell you the whole truth. We have just spent a thousand words saying so. It is still the right first move, because it answers the two questions that matter before any trend can exist: is the door unlocked (readiness), and is anyone walking through it (visibility).

Run a free scan at whocanfindme.com. No signup. Takes about ten seconds. Then run it again next week. The second reading is where measurement begins.

다음 글

GEO, Measurement

One question, ten hidden ones

When a buyer asks an AI one question, the engine quietly turns it into about ten. Seer Interactive measured 10.7 hidden sub-queries per prompt across 501 prompts on Gemini 3. Most sites answer the headline question and none of the ten. Here is how to find your gaps.

더 읽기
GEO

Different engines, different back doors

Claude searches the web through Brave. Gemini leans hard on YouTube. Perplexity is the most citation dense of the four. Where each engine gets its sources changes what you should do, and it is why we score every engine separately instead of blending an average.

더 읽기