
The most common objection to AI visibility measurement is also the most reasonable one: the answers keep changing. Ask the same question twice and you get two different results. So what exactly is being measured?
It is a fair challenge and it deserves a real answer rather than a reassurance.
The variance is worse than most people assume
Researchers at the University of St. Gallen examined how stable brand and source visibility actually is across ChatGPT, Gemini and Perplexity. Two findings matter here.
Cited sources showed only 34–42% overlap between consecutive days. Roughly 65% of the sources an engine cites change from one day to the next.Schulte, Bleeker & Kaufmann, University of St. Gallen, April 2026.
Identical prompts submitted minutes apart shared only about a third of their sources. Which means most of the instability is not the world changing. It is the model's own probabilistic nature.
Brand mentions were steadier than individual sources — Jaccard scores of 45–59% day to day — but still substantially variable.
So anyone who runs a single query, screenshots the result and presents it as a finding is showing you a coin toss. That includes anyone doing it to sell you something, and it includes the hotel marketing manager who checked once and concluded everything was fine.
Why this is a methodology problem, not a measurement problem
Variance of this kind is not unusual in measurement. It is the normal condition of anything sampled from a probabilistic system, and the discipline for handling it is well understood: fix everything you can, repeat, and report the trend rather than the reading.
Applied here, that means five things held constant.
- A locked query set. At least twenty-five questions per property, written in the language a guest would actually use, and never edited during an engagement. The moment a query changes, month two stops being comparable to month one.
- The same engines, every time. Four, measured identically. Because engines share as little as 13% of their cited sources, a change in engine mix would swamp any real movement.
- The same source market. Answers diverge substantially by location. A query about a Mumbai hotel run from London and from Mumbai are two different measurements, and mixing them produces a number that means nothing.
- A fresh session for every query. No chat history, no memory, no follow-up questions. One contaminated thread biases everything after it.
- An archived baseline. Measured and stored before anything is changed, and never overwritten. It is the fixed point the whole engagement is measured against.
Do that and the noise averages out across the set. A property that moves from 14% to 31% while its competitive set stays flat has achieved something specific and defensible. A property that moves from 14% to 17% has probably achieved nothing, and should be told so.
What a single reading is still good for
It is worth being fair to the one-off check. A single query is a poor measurement and a very good demonstration.
When a commercial director sees, in one screenshot, that an AI engine recommended four hotels in their market and named a competitor rather than them, that is not a statistic. It is a fact about a real conversation that a real guest could have had. It does not need to be repeatable to be worth knowing.
The mistake is treating that screenshot as a baseline. It is the thing that starts the conversation, not the thing that tracks progress.
How we handle it, and where we are still limited
We run twenty-five queries per property across four engines every month, from the client's actual source markets, and we archive every screenshot. That is a hundred measurements per property per month — enough to establish a stable trend without becoming unaffordable across a large portfolio. For groups above twenty-five properties we measure a representative sample and disclose the sampling method in the report.
Two honest limitations.
We do not run each query multiple times per day. The St. Gallen work suggests seven or eight runs per prompt for reliable brand-detection rates. At a hundred measurements per property per month, that would be eight hundred, which is not affordable at the price we charge. Our defence is breadth instead of depth — twenty-five different queries rather than one query eight times — and we think that is the better trade for a commercial engagement. We would rather say that than pretend the trade does not exist.
We report a trend, not a guarantee. A rising Share of Model means a property is being named more often across a fixed set of questions. It does not mean more bookings, and we do not claim it does.
The test to apply to anyone selling you this
Ask three questions.
How many queries, and are they locked? How many engines, and which? What country were they run from?
If the answers are vague, you are being sold a screenshot. If the answers are specific and published, you can audit the method before you buy it — which is the only reasonable basis on which to buy a measurement at all.