Editorial guide · Inspeccia
How to measure if ChatGPT recommends you: three practical methods for 2026
Talking about GEO without measuring is decoration. If you don't know how often ChatGPT mentions you when someone asks about your category, you're not doing generative SEO — you're writing content and crossing your fingers. There are three reasonable ways to stop guessing, and all of them take less time than you'd think.
I'll walk through each one below, with their honest pros and cons. After that, a word on when it makes sense to automate, and the mistakes we see in teams that get excited about their first measurement and blow it.
Why measurement matters more in GEO than in classical SEO
In classical SEO it's easy to know where you stand: open Search Console or Ahrefs and see your position per keyword. GEO has no equivalent dashboard. The question "does ChatGPT recommend me?" has no clean numeric answer because it depends on the prompt, the model, whether retrieval is on, and how the wind was blowing that day.
What you can do is build your own system. Repeatable, simple, anchored on the real questions your ideal customer asks an AI. Without it, any GEO investment is blind.
Method 1: manual prompt testing
The most basic, the most reliable, and the most underrated. The idea is to define ten to twenty prompts that represent how your potential customer would ask an LLM about your category, and run them against the models that matter for your market.
How it looks in practice:
- List the prompts. These aren't keywords, they're real questions. If you sell web audit tools: "what tool should I use to audit my website," "how do I check if my SEO is healthy," "cheaper Semrush alternative," "who does SEO audits for B2B SaaS." Think like your customer, not like an SEO.
- Run each prompt in ChatGPT (with retrieval enabled), Perplexity, Claude and Gemini. Those four cover roughly 95% of the generative search market right now.
- Log it in a spreadsheet: does your brand appear? In what position? In what tone? Does it get confused with a competitor? Does it say anything inaccurate?
- Repeat monthly. It's the month-over-month comparison that teaches, not the snapshot.
What takes time isn't running the prompts, it's deciding which ones to run. A solid list of fifteen representative prompts is worth more than a script blasting a hundred generic ones.
Useful variant: add five competitor prompts to the list. "Who is the main competitor to Semrush," "alternatives to Ubersuggest." If you show up there, you found qualified traffic you didn't know existed.
Method 2: brand mention monitoring
The second method looks sideways rather than forward. The premise is that LLMs learn from what gets published on the open web — Reddit, blogs, forums, press. The more your brand gets mentioned in those places consistently, the more reasons the model has to mention you in the future.
How to measure:
- Google Alerts, free, set up with your brand name and variants. You get an email every time you're mentioned on an indexed page.
- Mention or Brand24 if you have budget. They cover Reddit, Twitter/X, forums and comments better. Run thirty to a hundred dollars a month.
- Manual Reddit search: once a month, walk through the subreddits relevant to your category and search for your brand. Tedious but clear.
What you'll learn here is different from method 1. Method 1 tells you whether the AI mentions you today. Method 2 tells you whether you'll be mentioned six months from now, because today's Reddit mentions are what'll be in the next model's training corpus.
Method 3: live Brand SERP analysis
The fastest and the most underrated. Search your brand name on Google and look at the top ten results. That's what the AI sees the first time it talks about you.
What a healthy Brand SERP looks like:
- Position #1: your own site.
- Positions #2-5: official profiles (company LinkedIn, founder About, team page if you have one).
- Positions #6-10: authentic reviews on relevant vertical sites (G2, Capterra, Product Hunt depending on your niche), positive press if there's been any, quality content mentioning your brand.
What shouldn't be there:
- Automated scrapers copying your copy.
- Stale mixed reviews on generic directories that aren't relevant anymore.
- Forum threads with unanswered complaints.
- "Alternative to [your brand]" results above your own.
Write down what's wrong and put together a plan to push each problematic result out. It's among the best effort-to-impact ratios in all of GEO.
Running all three methods by hand is realistic for a brand with five to ten prompts and one model. If you want to scale — more prompts, several models, comparison against competitors, monthly cadence — you need automation.
That's exactly what Inspeccia does in every analysis: we run three real questions against ChatGPT about your industry and report how it describes you, whether it mentions you, which competitors it positions you against and what tone it uses. Plus live SEO data and a step-by-step plan. Start a free analysis or request a custom audit with your own prompts.
When automation actually pays off
If you're measuring a single brand with five prompts, do it by hand. The time cost is an hour a month, and the benefit of seeing the results with your own eyes is real. If you cross fifteen prompts, or want to cover three models seriously, the manual spreadsheet becomes painful by month two.
What's worth automating is the runs (scripts against the LLM APIs) and mention monitoring. What stays manual and valuable: deciding which prompts to use, reading responses with critical judgment, deciding what to do with the data.
Common mistakes
The most common is measuring prompts nobody would ever actually ask. Your customer doesn't ask ChatGPT "what's the best SaaS web audit solution on the market for mid-sized companies." They ask "what tool should I use to audit my site." If your prompts sound like a marketing brief, you're measuring something that doesn't exist.
The second is measuring once and declaring victory or defeat. A single data point isn't a trend. You need at least three monthly measurements before drawing conclusions, because LLMs have variance and the corpus changes.
The third is obsessing over one model. If you only measure ChatGPT, you miss Perplexity (increasingly used by marketers), Gemini (weighty because of its Google Search integration), and Claude (carrying more and more professional usage).
The fourth is not writing anything down. Measuring without recording is wasted work. A plain spreadsheet with date, prompt, model and result is worth more than any expensive tool you don't use regularly.
Frequently asked questions
How many prompts do I need to run to get a reliable read?
For a small brand with a clear vertical focus, ten to fifteen representative prompts are enough to spot trends. If your category is broader or you want to segment by buyer persona, push to thirty or fifty. What matters more than volume is consistency: use the same prompts month over month so you can compare.
Why does ChatGPT answer differently every time? How do I control that?
LLMs have temperature: a parameter that gives them some randomness so responses aren't identical. For measurement, run each prompt three to five times and look at the dominant response. If you're mentioned in four out of five runs, score it as one; if you appear in one out of five, score it as zero and flag it as unstable.
Should I use the ChatGPT API instead of the web interface?
Yes, especially if you want to scale. The web interface may have retrieval enabled (live web search) while the API defaults to no retrieval. For consistent measurement, pin the model version (gpt-4, gpt-4-turbo, etc.) and decide explicitly whether you want retrieval on or off. Most people want it on, because that's what real users experience.
From measuring to acting
Inspeccia's standard analysis includes exactly this measurement — three real questions to ChatGPT, live, with results captured — plus everything you need to start acting.