Guide · Inspeccia
AI visibility audit: what it actually measures (and what it cannot)
An AI visibility audit answers a question no dashboard answers: when a customer asks ChatGPT or Gemini about what you sell, do you show up? You find out by asking those questions again and recording four things — whether the models can reach your site, whether they name you, with which facts, and citing which sources. Everything else sold under that label is decoration.
Start with the uncomfortable part, because it explains why this exists as a category at all. There is no report to read.
Google says so in writing in its documentation for site owners: pages that appear in AI features — AI Overviews and AI Mode — are "included in the overall search traffic in Search Console". They are in there, mixed in with the rest of Search, with no row of their own to filter. You can see that you got impressions; you cannot see how many of them happened inside a generated answer.
On the other side there is not even that. OpenAI publishes no panel where a site owner can see how many times ChatGPT named their company. Neither does Perplexity, neither does Claude. The surface where your customers are asking is, for you, opaque by design.
The audit exists to fill that gap. If nobody hands you the number, the only way to know it is to rebuild it: ask the questions, read the answers. It sounds rudimentary because it is. It is also the only thing there is.
The four layers a serious audit measures
The word "visibility" hides four different problems, with different fixes and very different costs. An audit that does not separate them leaves you with a number and no plan. Here they are, in the order worth checking them.
1. Access: can they reach you at all?
This is the boring layer and the one most often broken. Before asking whether a model recommends you, ask whether it can read you.
OpenAI documents that OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features" — that is, it is the one that makes you eligible to appear cited — and that GPTBot is a different crawler, the one tied to foundation models. Two separate decisions, and plenty of people block them together by accident: they wanted out of training and deleted themselves from search as well.
On Google's side, the mirror image is snippet control. Its documentation points to nosnippet, data-nosnippet, max-snippet and noindex as the way to limit what is shown from your pages. They work, which is exactly why they hurt when they are set too tightly: a short max-snippet inherited from another era leaves you indexed and, at the same time, with nothing quotable.
This layer is audited with a technical review, not by asking any model anything. And it is the only one of the four where a finding gets fixed the same day.
2. Mention: do they name you when nobody asks about you?
Here is the most common mistake people make auditing themselves: asking ChatGPT about their own brand. That does not measure visibility, it measures memory. Give it the name and the model will talk about you; the customer who matters does not know your name yet.
The question that counts is the customer's: by category, by problem, by area. "Who does X in Y?", "what is the best option for Z?". And the useful answer is not yes or no, it is how many times out of how many. Repeat the same question across separate sessions and you will see yourself appear once and vanish twice. That share is the metric; a single run's yes or no is nothing.
3. Accuracy: with which facts do they name you?
Showing up badly described can cost more than not showing up. A model that names you with an old price, a service you dropped or the wrong city is doing sales work against you, with the authority that comes from sitting inside an answer rather than an ad.
This layer is audited by comparing what the model asserts against what is true, field by field: what you do, where, since when, at what price, for whom. When a gap appears, the fix is almost never in the model: it is in the source the model read.
4. Sources: who is it citing?
The layer most people skip and the one that explains the most. When a search-enabled model answers about your company, it opens specific pages. Knowing which ones tells you who is writing your reputation: your own site, a directory, a forum, a four-year-old article, a competitor's blog.
It is also the layer that turns an audit into a work plan. If the sources describing you are third-party and out of date, you already know where to intervene, and it is not your homepage.
Order matters. Auditing mention and accuracy before access wastes the effort: if OAI-SearchBot is blocked or your snippets are capped, the other three layers will come back bad for a reason that has nothing to do with your content. Always start at the door.
The three things no audit can promise you
This section is what separates a diagnosis from a brochure. If whoever audits you does not tell you these three, they are selling certainty they do not have.
There is no traffic number. As above: Google folds AI features into overall Search totals in Search Console. Anyone showing you "visits from AI Overviews" is estimating, and should be using that word.
The same question does not give the same answer. Generated answers vary between sessions, between countries and between model versions. An audit that runs each question once is not measuring your visibility: it is photographing something that moves. Repetition is not a methodological luxury here, it is the methodology.
One user question is not one search. Google describes AI Mode as running a technique it calls query fan-out: "breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf". In Deep Search, it adds, the system "can issue hundreds of searches". So behind the one question you audit there are dozens of queries you never see, and they also decide whether you show up. Auditing by isolated keywords understates the problem by design.
How to tell an audit from a smoke report
A fashionable category attracts a lot of PDFs with a speedometer on page one. Four questions are enough to sort them.
- How many questions did it run, and how many times each? No repetitions, no measurement: just an anecdote.
- Does it show the raw answers? A score without the text that produced it is not verifiable. You should be able to read what the model actually said.
- Does it name the sources the model opened? If not, the actionable half of the work was not done.
- Does it tell you what it does not know? Variance, country, date, model. An audit that does not state its limits is hiding the margin of error, not removing it.
And one concrete red flag: if the report promises to get you into AI Overviews through special markup or new files, it contradicts the source. Google writes that "there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary", and that you "don't need to create new machine readable files, AI text files, or markup" to appear in those features. Anyone selling you that shortcut is selling something Google itself says does not exist.
Where to start this week
Before hiring anything, run the short version yourself. Pick the five questions a customer would use to find you — category and area, never your brand. Ask them in ChatGPT and in Gemini, three times each, in separate sessions. Record in a table: whether you appear, where in the answer, with which facts, and which sources get cited.
That one-hour table already tells you which of the four layers is failing, which is the expensive decision. If you never appear, check access. If you appear sometimes, it is authority. If you always appear but described wrong, it is a data and sources problem, and that is the fastest to fix. Whatever you buy afterwards — tooling, a service, anything — is scale on top of that same method, not a different one.
Frequently asked questions
What exactly is an AI visibility audit?
It is the work of reconstructing, question by question, what AI assistants say when someone searches for what you sell. It is not reading a report, because there is no report to read. Google documents that pages appearing in AI features are "included in the overall search traffic in Search Console", with no breakout of their own, and ChatGPT publishes no dashboard for site owners. So an audit asks the questions again and records four things: whether the models can reach your site, whether they name you, with which facts, and citing which sources.
How is this different from a normal SEO rank report?
There is no rank to report. In classic SEO there is an ordered, fairly stable ranking for a keyword; in a generated answer there is no list of ten blue links, there is a paragraph that either names you or does not. On top of that, the same question can produce different answers in two consecutive sessions. That is why an honest audit works with repetitions and with the share of runs in which you appear, rather than a single number pretending to be a position.
Can I run one myself without buying anything?
Yes, and you should start there. Open ChatGPT and Gemini, ask the way a customer would ask - by category and area, never by your brand name - and record whether you appear, where in the answer, and which sources are cited. Run each question three times in separate sessions to see the variance. What a tool buys you is not magic: it is scale and repetition, which is exactly what becomes unworkable by hand once you go past ten questions.
Is it worth it for a small or local business?
More so, not less. A small business usually has the cheapest version of the problem: few external sources describing it and inconsistent details across its site, its listing and its profiles. That kind of error gets fixed in weeks. The audit tells you whether your problem is access (models cannot read you), data (they read you and get it wrong) or authority (they read you correctly and still pick someone else), and each has a different fix.
Why do two AI visibility tools give me different numbers?
Because they measure different things and none of them measures "the truth". They differ in which questions they run, how many times they repeat them, in which country and against which model. Since answers vary between sessions, a tool that runs a handful of prompts once a month and one that runs many every week cannot agree. The useful question is not which one is right, but which one shows you the variance instead of hiding it behind a round number.
Sources cited
- Google Search Central — "AI features and your website" (sites appearing in AI features are included in the overall search traffic in Search Console, within the Web search type, with no separate breakout; no additional requirements or special optimizations to appear in AI Overviews or AI Mode; no new machine readable files, AI text files or markup needed; snippet controls
nosnippet,data-nosnippet,max-snippetandnoindex). Official documentation. - OpenAI — "Overview of OpenAI Crawlers" (
OAI-SearchBotis used to surface websites in search results in ChatGPT's search features;GPTBotrelates to foundation models;ChatGPT-Userto user-initiated actions). Official documentation. - Google — "AI Mode in Search" (query fan-out technique: breaking the question into subtopics and issuing a multitude of queries simultaneously; Deep Search can issue hundreds of searches). Published 20 May 2025. Official announcement.
The automated version of these four layers
Our analysis does exactly this: it checks whether AI crawlers can read you, asks ChatGPT and Gemini live about your category, compares what they say about your company against what is true, and shows you which sources they opened. Nothing to install.