Guide · Inspeccia

Why every AI visibility tool gives you a different score

An agency opens two free trials on the same Tuesday, for the same client, a twelve-person accounting firm. One tool reports the brand shows up in 8% of answers. The other says 40%. Neither explains where the number came from. The uncomfortable part is that both may be perfectly well built, because the trouble starts earlier: the same question, put to the same model twice in a row, does not come back the same.

That is not a hallway impression. SE Ranking ran the same 10,000 keywords through Google's AI Mode three times on a single day and compared the links returned. Across 122,617 links, the average overlap of exact URLs between the three passes was 9.2%. On 2,006 keywords — 21.2% of the set — not one URL matched. Six keywords out of ten thousand matched completely.

When the substrate you are measuring moves like that, any score built on top inherits the movement. This guide explains why it happens, what we found when we repeated our own audits, and how to work with a number that wobbles without giving up on deciding things with it.

Almost nobody publishes how the score is calculated

Try a short exercise. Open the pages of five AI visibility tools and look for how many times each prompt is repeated before averaging, against which exact model, from which country, and what counts as "appearing" — a brand mention, a link, a link in first position. You will find the score in the first screenshot and the method nowhere.

The industry admits the hole when it measures itself. In its 2026 index, built on 126 million prompts from January through April, Semrush found that 45% of marketing leaders cannot accurately measure their brand's visibility inside AI-generated answers, and that only 9% have the tools to track every relevant metric across platforms. Nine in a hundred.

There is a widespread assumption worth dismantling here: that an AI visibility score behaves like a Google position. An organic ranking is stable enough that two people checking it on the same day see roughly the same thing, which is why an entire reporting industry could be built around it. The SE Ranking figure says that on generative surfaces the assumption collapses, even within one day and one engine. A score that does not say how many times it was measured is not a data point. It is an anecdote with two decimal places.

Five reasons the same question gives different answers

1. Every tool asks something else

No tool measures you against "AI". It measures you against a list of questions somebody chose, and that list is half the result. A panel asking "best accounting firms in Austin" and one asking "who can handle the books for a small business" are measuring different markets under the same label.

Conductor quantified this usefully: 14,000 API calls, ten industries by seven intent types by four engines, fifty runs per cell. The finding was not that one industry is more volatile than another, but that intent type predicts consistency better than industry does. And the worst intent is purchase. Out of every ten unique brands that appeared across two runs of the same purchase prompt, only four showed up in both. The moment you most want to be named is the moment that repeats least.

2. The model shifts underneath you while you measure

A visibility dashboard is aimed at a target that updates without notice. Writesonic ran 631,999 prompt-model pairs across seven platforms between March and June 2026, repeating each prompt at least ten times, and measured how much the top spot rotates: on ChatGPT, 52% of number-one brands change on the same prompt. In the same body of work, comparing which sources ChatGPT, Gemini, Perplexity and AI Overviews cite for identical prompts, only 3.8% of sources appear on all four, and between 72% and 73% of cited domains appear on exactly one engine and nowhere else.

One consequence follows that almost no report respects: any citation percentage is a photograph with a date on it. If your August deck quotes a May figure without saying when it was taken, it is already lying a little.

3. The person asking is part of the question

OpenAI documents that ChatGPT rewrites the user's prompt before sending it to search, and that when memory is on it may fold what it knows about that person into the rewrite. Someone who has spent months discussing bookkeeping for freelancers does not generate the same query as someone opening the conversation cold, even if they both type the same sentence.

That breaks the comparability of any measurement taken from a real account. Measure inside your own ChatGPT, with your history and your saved projects, and you are measuring your bubble. The baseline method — clean window, logged out, memory off — is in the guide to measuring your AI visibility, and it is the step most people skip.

4. Where you ask from changes who shows up

General location derived from the IP address feeds into the rewritten query, which means two measurement servers in two countries cannot return the same result even if they wanted to. In our own audits the gap by market is large: among domains with at least thirty cases, 97% of US sites drew no mention at all when we asked about their category, against 51% of the ones based in Cape Town.

That contrast says less about geography than about competition: where thousands of rivals share a category, the odds of being named in a short answer collapse. The operational consequence stands either way. If your tool queries from Virginia and your client sells in Montevideo, the number on screen is not describing your client's market.

5. The model is not deterministic with itself either

The most uncomfortable reason is last, because no measurement methodology fixes it. In September 2025, Thinking Machines Lab published an analysis of why inference endpoints return different outputs even at temperature zero, and the explanation is not the one usually repeated. The primary cause, they write, is that server load — and therefore the size of the batch of requests processed together — varies nondeterministically. The compute kernels are not batch-invariant, so the result depends on how many other people happened to be asking at that instant.

"From the perspective of an individual user, the other concurrent users are not an 'input' to the system but rather a nondeterministic property of the system."

It is a solvable problem: the same paper ships batch-invariant kernels that produce reproducible inference. But that gets applied by whoever serves the model, not by the tool querying it, and no consumer chat product publishes whether it does so.

What we found when we repeated our own audits

Between May 14 and July 28, 2026, we ran 2,324 audits across 2,021 unique domains. Each one puts real questions to a model with live web search and records which sources it consulted. On 167 domains the analysis was run two or more times, averaging 2.4 runs, which allows the cross-tab nobody publishes: what happens when you measure the same thing twice.

Among the 51 domains that were visible at least once, 56.9% produced a different result on repeat, with an average swing of 30.8 points of mention rate. More than half of the brands that ever appear change position when you ask again, and the typical jump is thirty points. That is the real size of the noise.

Now the honest caveat, because the figure above is easy to inflate. Count all 167 domains and the average swing drops to 9.4 points, with only 17.4% moving by a third or more. That sounds far steadier. It is not: most of those domains were invisible on both runs, and a repeated zero does not measure stability, it measures absence. Averaging the invisible ones into the pool is the easiest way to publish a reassuring headline on data that says something else.

The limits of the cross-tab, out loud: a single LLM provider, a cap of twelve source domains stored per analysis, self-declared categories, and repeats that happened whenever a user decided to re-run rather than on a controlled schedule. This is not a laboratory experiment. It is what production looks like, which is where your report lives.

An honest AI visibility number always arrives with three things attached: the exact question, the date, and the number of runs. Missing any of the three, it is not comparable to anything, including its own past self.

At Inspeccia we put real questions to a model with live search and show you what comes back: whether it names you, who it names instead, and which sources it cites to describe you. Start a free analysis.

How many runs a number needs

What follows is a rule of craft derived from the observed swing, not a theorem: nobody has published a confidence interval for AI visibility, and anyone selling you a magic number of runs as established statistics is decorating. The reasoning is plain. If the typical swing between two runs of a visible domain is around thirty points, one run cannot separate a 20 from a 50. You need enough repeats that the average stops being dominated by a single throw.

  1. Freeze the prompt set before you measure. Write it down, date it, store it. Edit the questions after seeing the result and what you are measuring is your own editing. This step has produced more attractive reports and fewer good decisions than any other.
  2. Repeat each question three to five times, on different days. Three runs spread across days capture the variation better than ten back-to-back in one afternoon, because the model version and the search index change between days, not between minutes.
  3. Always measure in a clean session. Private window, logged out, memory off, no saved projects. And from a location that resembles the client's market rather than your office.
  4. Report a range and a median, never a bare score. "Between 20% and 55%, median 35%, across five runs from August 3 to 10" is a sentence that survives being questioned. "Visibility: 35" is not.
  5. Set an action threshold. If the movement between two checkpoints is smaller than the swing you have already observed, do not count it as change. Celebrating noise is worse than not measuring, because it makes you repeat the wrong action.
  6. Do not compare scores across tools. There is no shared definition of "AI share of voice" and no public benchmark prompt set. Two separately invented scales do not subtract. Compare your number against your own previous number, same list, same method.

And always pair it with something countable. Referral traffic from ChatGPT, Perplexity or Gemini lands in your analytics and does not depend on the model's mood; isolating it is covered in the Google Analytics 4 guide. A dashboard swinging thirty points and a session counter creeping upward are telling the same story from two sides, and the second one is what the client signs off on.

What to send a client every month

For an agency this stops being a methodology debate and becomes the problem of what goes in the August PDF. The escape hatch people reach for too often — three screenshots of ChatGPT naming the client — has a fatal defect. It is irreproducible. The client opens ChatGPT that same afternoon, sees something else, and the meeting turns into an argument about whether the screenshot was real.

A defensible report changes the object. Instead of promising a score that goes up, it hands over evidence that can be lifted again:

  • The verbatim prompt list, with date and run count. Without it, nothing else is checkable.
  • This month's range next to last month's, with the action threshold marked.
  • Who appears when the client does not. This is the section that gets read and the one almost nobody includes: the roster of competitors occupying the answer is more actionable than any score.
  • Which sources the model cited to describe the brand. In our audits it cites 2.65 sources on average, and in 26.9% of cases the client's own site is not among them. When that happens, the month's work argues for itself.
  • AI referral traffic, the one number in the folder that holds still.

It is worth checking the mirror before selling any of this. Among the digital marketing agencies that have run through Inspeccia — thirty cases, a small sample and it should be said — 90% drew no mention when we asked about their own category: twenty-seven out of thirty. The industry that sells AI visibility does not have it. Poor sales pitch, excellent starting point for a case study of your own, which is the angle we develop in the guide to AI visibility for agencies.

What an honest number can promise

Measuring AI visibility is still worth doing. Treating it like a rank position is what fails, because the object does not behave that way and no tool can fix that from the outside. What survives repetition is the structure of the answer rather than the score: whether you get named, with which attributes, citing whom.

That layer is the one you can move with work, and it is also the one that explains the score when the score moves. If you want the full conceptual frame for why optimising for generative answers plays by its own rules, it is in the introduction to GEO. The rest is saying how many times you measured, and saying it before anyone asks.

Frequently asked questions

Why do two tools report such different scores for the same brand?

Because they are almost never measuring the same thing. Each one uses its own prompt set, its own mix of engines, its own outbound location and its own definition of "appearing": for some a brand mention counts, for others only a link does. On top of that, the model's answer moves between runs even when nothing else changes. Conductor measured 14,000 API calls and found that on purchase-intent prompts, out of every ten brands that showed up across two runs, only four appeared in both. Two honest tools can land far apart without either being broken.

How many times do you have to repeat a measurement for it to be usable?

There is no canonical number, and be wary of anyone who hands you one as settled science. What the data does support is a rule of craft: when the observed swing between two runs of the same domain sits around thirty points, a single run cannot tell a 20 from a 50. Three to five runs of the same prompt, spread across different days and taken in logged-out sessions, give you a range you can work with. If the range comes back wide, the finding is that it is wide.

Can I compare my score from one tool against another tool's score?

No, and this is the most common trap in reporting meetings. Visibility scores are not standardised: there is no shared definition of "AI share of voice", no public benchmark prompt set, and no convention on how many runs get averaged. Comparing two vendors' scores means comparing two scales invented separately. What does work is comparing your number against your own previous number, using the same prompt list, the same engine and the same counting method every time.

If the number moves that much, is measuring worth anything?

It is, just not for the thing the marketing promises. A bare score cannot carry a decision. What can carry one is everything else that comes out of the same measurement: whether the model names you or names someone else, what it attributes to you, which sources it cites when describing you, and whether those sources are yours or other people's. That layer moves far less between runs and is directly actionable. In our audits the model cites 2.65 sources on average per brand, and in 26.9% of cases the brand's own site is not among them.

What should a monthly client report actually contain?

Nothing that depends on a screenshot. A defensible report has five parts: the verbatim prompt list with its date and run count; a range instead of a single score; who appears when the client does not; which sources the model cited to describe the brand; and AI referral traffic from analytics, the one number in the folder that does not depend on the model's mood. If this month's movement is smaller than the normal swing between runs, saying so is also reporting.

Sources cited

  1. SE Ranking — "AI Mode Research": the same 10,000 keywords parsed three times in one day across 122,617 links; average exact-URL overlap of 9.2%; 21.2% of keywords with no URL in common. Study.
  2. Conductor — "AI Brand Recommendation Study": 14,000 API calls (10 industries × 7 intent types × 4 engines × 50 runs); purchase-intent prompts are the least consistent, with four of every ten brands appearing in both of two runs. Analysis.
  3. Writesonic — "Do AI Engines Cite the Same Sources?": 161,286 prompts for the source-overlap study (3.8% of sources universal across four engines; 72–73% of cited domains on a single engine) plus a ranking stability study over 631,999 prompt-model pairs in which 52% of ChatGPT's number-one positions rotate on the same prompt. Study.
  4. Semrush — AI Visibility Index 2026, 126 million prompts from January to April 2026: 45% of marketing leaders cannot accurately measure brand visibility in AI answers and only 9% have tools to track all relevant metrics. Announcement.
  5. Thinking Machines Lab — "Defeating Nondeterminism in LLM Inference", September 10, 2025: nondeterministic variation in server load and batch size explains why inference is not reproducible even at temperature zero; proposes batch-invariant kernels. Article.
  6. OpenAI Help Center — "ChatGPT Search": rewriting of the user's prompt into a search query using general location derived from the IP address, and memory's influence on that rewrite. Help article.
  7. Inspeccia — analyses_v2 database: 2,324 audits across 2,021 unique domains between May 14 and July 28, 2026; stability cross-tab over 167 domains with two or more runs; 2.65 sources cited on average per brand; 26.9% of brands with no citation of their own site; 27 of 30 digital marketing agencies with no mention in their category.

Look at the evidence before the score

See what AI answers about your brand today: whether it names you, who shows up instead, and which sources it cites to describe you. The analysis is free and takes a few minutes.