Honest measurement · August 31, 2026 · 5 min read
What AI Visibility Tools Actually Measure
Every AI visibility tool works by sampling the assistants with synthetic prompts. Here is what that measures, what it cannot measure, and which claims deserve skepticism.
We build an AI visibility tool, so read this with that in mind. But precisely because we build one, we can tell you how the sausage is made: every tool in this category, ours included, works by asking the assistants questions and analyzing the answers. That is the whole trick. The differences between tools are in how well they do it and how honestly they describe it.
This matters because the category’s marketing frequently implies something else: that these tools can see what real users are asking ChatGPT and what it tells them. They cannot. Nobody can.
The two kinds of data, and the only two kinds
Synthetic sampling. The tool maintains a set of prompts (“best CRM for startups”), runs them against the assistants on a schedule via their APIs, and parses the answers for brand mentions, positions, and citations. This is polling: structured, repeatable, and entirely under the tool’s control. It answers the question “when the assistants are asked X, what do they tend to say?”
Referral analytics. Your own website analytics can attribute sessions to chatgpt.com or perplexity.ai referrers. This is real user behavior, but it only captures the moment someone clicks through to your site from an assistant, which most conversations never produce. It answers “how many people arrived here from an assistant?”, not “what did the assistant say about me?”
Everything an AI visibility tool shows you is built from one or both of these. There is no third source. Conversations between users and ChatGPT are private; OpenAI does not sell them, and no vendor has a side door.
What synthetic sampling genuinely tells you
Used properly, sampling is a good instrument. It gives you:
- Tendencies with confidence. Ask a prompt 300 times over a month and you know, with real statistical footing, how often each brand appears and in what position. This is how the core metrics are computed.
- Trend lines. When a model update ships or your content strategy lands, sampled metrics move. The direction and timing of that movement is genuinely informative.
- Citation intelligence. Sampled answers reveal which sources assistants lean on in your category. That list is the closest thing GEO has to a ranking factor report, and it converts directly into a to-do list. Our guide to checking your brand in ChatGPT shows how to read it.
- Competitive position. Sampling treats you and your competitors identically, so share-of-voice comparisons are fair even though the absolute numbers are synthetic.
What it cannot tell you, no matter the vendor
- What real users asked. Your prompt set is a hypothesis about buyer questions. If real buyers phrase things differently, your samples measure the wrong conversation. Good tools let you evolve the prompt set; no tool can confirm it matches reality.
- How many people heard any answer. A 40% visibility rate does not mean 40% of your market hears about you. Sampling has no denominator in humans; it measures the assistant, not the audience.
- Personalized answers. Real users have memory, custom instructions, and conversation history. API sampling is clean-room by design. The assistant your prospect talks to is subtly different from the one any tool measures.
- Revenue impact, directly. The path from “mentioned in answers” to “signed up” runs through territory nobody can instrument end to end. Correlating visibility trends with your own referral and signup data is the honest approximation, and it is an approximation.
Questions that separate honest tools from the rest
If you are evaluating tools in this category (there are several: Profound, Peec, Otterly, us; we compare them side by side in CitePrism vs Profound), these questions cut through the marketing:
- “Can I read the raw answers behind this chart?” If the tool will not show you the underlying responses, you cannot audit anything it claims.
- “How many samples is this number based on?” Any percentage without a sample size attached is decoration. Small n should be visibly flagged, not hidden.
- “What happens when I change my prompt set?” Changing questions changes numbers. A trustworthy tool marks the change on the timeline instead of letting the discontinuity masquerade as a trend.
- “Where does this data actually come from?” If the answer implies access to real user conversations, walk away. You now know why.
CitePrism’s answers: every chart clicks through to stored raw responses, low-sample metrics are greyed out until they cross a minimum, every metric definition is written out on the site, and each run shows what it cost. You can start free on your own model keys. We built it this way because measurement tools earn trust with auditability, not adjectives.
FAQ
So are AI visibility tools worth using at all?
Yes, for the same reason polls are worth running despite not being elections. Sampling gives you tendencies, trends, and citation sources you cannot get any other way, and those are enough to steer a GEO strategy. The failure mode is not using a sampling tool; it is believing you bought a wiretap.
What is the difference between AI visibility tracking and referral traffic reports?
Visibility tracking samples the assistants to see what they say about you. Referral reports count users who clicked through to your site from an assistant. The first measures the answer, the second measures a click after the answer. They are complementary, and neither one is “real AI search analytics” on its own.
Why do numbers differ between two AI visibility tools?
Different prompt sets, different sample sizes, different schedules, different mention-detection logic, sometimes different model versions. Both can be internally consistent and still disagree. Compare trends within one tool rather than absolute numbers across tools.
Does a high visibility rate mean customers are hearing about us?
Not directly. It means the assistants tend to mention you when asked the questions in the prompt set. Whether many humans ask those questions, and whether they act on the answers, is outside what any tool can observe. Pair visibility trends with your own referral and signup data before drawing revenue conclusions.