In this article
I've personally tested several of the tools that promise to tell you exactly how visible you are in ChatGPT, Gemini, Claude, and the rest. For a while, I used PromptWatch, was actually quite happy with the experience, and yet ended up trusting the numbers less and less.
It may well be that I'm generally too sceptical here. But after digging into how the whole category actually works, I believe my scepticism was justified, for reasons I hadn't quite articulated at the time.
This is the fourth part of an ongoing series. Part 1 is about how AI models actually search and select sources, and part 5 is about GA4 setup. This one is about something else entirely: the tools that are supposed to tell you where you stand, right now.
What these tools actually do
The principle behind almost all tools in this category is the same. You build a library of questions, preferably based on what your customers actually ask. The tool sends these questions to ChatGPT, Gemini, Claude, Perplexity, and other models on a fixed schedule, and records if and how your brand appears in the answers.
What they report on is usually a combination of a few metrics:
- How often you are mentioned at all.
- Where you stand compared to your competitors on the same questions.
- Which sources actually drive mentions, such as a Reddit thread, a news article, or a YouTube video.
- Where there are gaps, i.e., questions where the AI answers well within your category but doesn't mention you at all.
It sounds neat on paper. The problem lies in what happens between the query and the answer.
One question is never the same question twice
Before I move on to what the various tools actually claim to deliver, it's worth pausing on something more fundamental. How can a tool even claim to measure "your visibility in ChatGPT" when every single person who actually uses ChatGPT phrases their question slightly differently?
This isn't just a theoretical problem. Two people asking for “best accountant in Oslo” and “accountant Oslo recommendation” might get different answers, with different sources cited, even if the intent behind the questions is identical.
The model itself also introduces variation; the same person asking the same question twice in a row could, in principle, get two different answers. This is because these models, unlike a traditional Google search, don't always give the exact same answer to the exact same question every single time.

What these tools actually do is not measure your visibility.
They measure how a fixed, limited set of questions are answered at a given time and use that as an approximation of the real picture.
The quality of that approximation depends entirely on how well the prompt library actually reflects what real customers ask, and how many times each question is run to smooth out random variation. A tool that runs fifty fixed questions once a week gives you a completely different level of precision than one that runs several thousand variations daily, even if both present the result as a single, neat percentage.
This is also why I mention server logs several times in this series. A server log tells you something completely different from a synthetic prompt experiment.
It shows you that a named AI bot, for example PerplexityBot or OAI-SearchBot, actually visited a specific page on your site at a specific time because a real person's real question somewhere triggered that visit. That's data from reality, not a simulation of it.
The problem is that a log only shows you that your page was considered, not whether it actually ended up being cited in the given answer. So the two methods answer different questions, and a good tool should ideally combine both instead of relying on just one.
The numbers fluctuate more than you think
Here's something I wasn't aware of until I dug into it. One analysis found that between 40 and 60 per cent of the domains cited by the major AI platforms change from month to month.
This means that a tool showing you a visibility percentage today might show you a completely different number in four weeks, not because something you did changed, but because the models themselves are moving targets.
How the different tools claim to measure visibility
I've looked through the available comparisons, and a recurring finding is worth mentioning right away: almost all of these comparisons are written by a vendor in the category, who naturally presents their own solution in the best possible light. With that disclaimer clearly stated, here's what they actually claim to do differently.
Profound is the category leader in terms of funding, having raised over 150 million dollars. It tracks in real-time across more than ten engines and has a dedicated feature for analysing which AI bots are actually visiting your website via a Cloudflare integration. The price starts at around 99 dollars a month, but a realistic entry point for something useful is closer to 399, and even then, you only get coverage for three engines.
Peec AI has focused on what they call the most accurate monitoring on the market, with simulations of real user interactions and reporting in over 115 languages. They offer prompt volume data—i.e., how many people are actually asking a given question—which most competitors do not. The price starts at around 89 euros a month.
Otterly.ai is the most budget-friendly option, starting from around 29 dollars a month, with Semrush integration being a clear advantage for those already using that tool.
PromptWatch, the one I have the most experience with, markets itself with the broadest publicly stated number of engines in the category, including analysis of AI crawler logs and a content agent supposedly driven by actual citation data. But an independent review from August 2026 points to something I recognise: the pricing ladder goes from 95 to 579 dollars, and even on the more expensive plans, there's a cap of four tracked models at a time, despite ten model logos being displayed on the pricing page. The review concludes with 4 out of 5 but explicitly notes that no such site can determine if the visibility numbers are actually stable from week to week, which they themselves call the most important criterion in the entire category.
That last point is the core of my scepticism. No matter how good the feature set is, none of the vendors say anything specific about how much noise is present in their own figures.

What I recommend
I'm not settling on a single "use this tool" answer, for the same reason that part 1 of this series concluded that there's no general roadmap for AI visibility. But here is some concrete advice, based on what I've seen:
-
Look for engine coverage that actually matches your audience, not the highest number on the pricing page. Several of the tools list ten models in their marketing but actually lock usage to three or four on the plans people actually buy.
-
Ask specifically about prompt volume and update frequency. A tool that runs few questions infrequently gives you a high-variance snapshot, not a reliable trend. So far, Peec AI is the only one I've seen that actually provides prompt volume as a separate figure.
-
Treat any single measurement as noisy, not as the truth. Given that 40 to 60 per cent of cited domains change monthly regardless, you should look at the trend over several months, not react to a single snapshot. This is probably the most important adjustment I should have made myself earlier.
-
Be extra sceptical of case studies and comparisons from the vendor itself. Almost every “best tools of 2026” article I've read has been written by someone with a commercial interest in the outcome, often without it being clearly disclosed.
-
Don't let one tool's low numbers alone decide a major decision, which was exactly my own experience with PromptWatch. Low numbers might mean low actual visibility, but they could also mean the tool has narrower coverage, fewer prompts, or more noise than the competitor you're comparing it with.
For now, I'm settling on using this type of tool as a guideline and a trend to follow over time, not as a precise metric to steer by. Besides, there's a lot you can do with free tools before paying for anything in this category.
Why I'm leaning on keyword analysis in the meantime
What I'm focusing on more in the meantime is sticking to classic keyword analysis, simply because it's the only data in this landscape that we actually have good, mature documentation for. The idea is simple: a keyword with high monthly search volume should also give rise to many prompts, as it reflects something people are already wondering about.
This is true to a point, but not all the way, and it's worth knowing why. Google is still strongest when someone knows what they want and is ready to act.
AI models are increasingly used for something else: a sparring session before a decision has even been made. “I'm considering switching accountants, what should I think about?” is a typical AI prompt that would never be typed into Google.
Profound has called this a “whitespace”: topics with a high volume of AI questions but low or no traditional search volume. This means that keyword volume as a proxy works best for purchase-intent, transactional questions and systematically underestimates the prompt volume for the more exploratory, advisory use that AI models are often preferred for.
This isn't to say that the visibility tools are worthless. It's to say that the category is still young enough that you should keep a bit of distance from making decisions based on a single number from a single tool.





