How we measure AI visibility
Every competitor hides how their numbers are made. Here are ours, in full.
What we collect
We ask ChatGPT the questions real buyers ask, in its actual consumer interface, from your project's chosen country and language. Not an API shortcut: the same ChatGPT your customers use.
Each prompt runs seven times a week. We save the full answer, the model version, the timestamp, every source it cited, every web search it ran, and everything those searches returned.
Every number in your dashboard links back to those raw answers. You can always read them yourself.
Why seven times
Ask AI the same question twice and you get two different answers. So a single check tells you almost nothing.
Seven samples a week is the same monthly volume as tools that advertise "daily tracking". The difference is what we do with the samples: we treat them as a distribution, not a fact.
That is why every rate we show carries a 95% confidence interval. We use the Wilson score interval for single rates and a cluster bootstrap over prompts for aggregates.
The Visibility Score
One number, 0 to 100, for how visible a brand is across its tracked prompts. Here is the exact formula, nothing hidden:
Score = 100 x sum(w x rate x position) / sum(w), summed over prompts.
rate is the simple part: the share of that week's answers that mentioned the brand.
position rewards where you appear. It is 0.5 + 0.5 x (average prominence / 100), so a buried mention still counts half, and leading the answer counts double a buried one.
w weights each prompt by how much people actually ask about its topic: 1 + log10 of the topic's modeled monthly demand (floor 10), and just 1 when the topic has no modeled demand.
Prompts weigh equally inside a topic, so demand works at topic level. Big topics count for more, and tracking lots of prompts in one topic does not inflate anything by itself.
The logarithm is there so one giant topic cannot run the whole score. It also keeps the score steady when the volume model is off.
A worked example, two prompts in different topics.
Prompt A's topic is modeled at 1,000 asks a month, so w = 4. The brand shows up in 6 of 7 answers (rate 0.857) at prominence 80 (position 0.9).
Prompt B's topic has no modeled demand, so w = 1. The brand shows up in 1 of 7 (rate 0.143) at prominence 20 (position 0.6).
Score = 100 x (4 x 0.857 x 0.9 + 1 x 0.143 x 0.6) / (4 + 1) = 63.
Competitors are scored with the identical formula and weights, so ranks are comparable. Confidence intervals come from the same cluster bootstrap as every other aggregate.
We expect to refine the formula as real multi-market data accumulates. Any change gets announced here and recomputed consistently across history, so trends never silently break.
How we decide something changed
Most tools draw a green arrow the moment a number ticks up. We do not.
A change only gets flagged when it passes a significance test: Fisher's exact test per prompt, a bootstrap difference test for aggregates.
At seven samples a week, most weeks genuinely show no significant change. When that happens, we say so plainly instead of decorating noise.
The same gate applies to the movement icons on cited sources.
Mentions, prominence, sentiment
A mention is your brand named anywhere in an answer. We match on your names and aliases, with a language model resolving ambiguous names in context.
Prominence measures where you appear: first, in a ranked list, or in passing. It is a published 0 to 100 formula.
Sentiment is a language-model read of how each mention characterizes the brand: positive, neutral or negative. We report the share of positive mentions, with a Wilson interval like everything else.
Naive LLM sentiment scores are unreliable, which is exactly why ours ships with two guards.
First, the score holds until 5 classified mentions in a week. Below that, one flipped mention would move it by 25 points or more.
Second, you always see the full positive/neutral/negative split, never a single averaged number.
When a week sits under the floor but the brand had a measured score within the last 4 weeks, we show that last measured score, clearly marked as carried. The trend line carries the last measured value forward the same way.
A brand with no measured week inside that lookback shows no score at all. We would rather show a gap than an invention.
Alongside the score, we publish the exact descriptive wording the answers used ("affordable", "steep learning curve"), extracted per mention and counted. So you see what is actually being said, not just how positive it sounds.
Topics
Your prompts are grouped into topics. A language model suggests the grouping at setup, and you can change it any time.
Every metric is recomputed per topic from the same per-prompt data. That is why topic scores work retroactively across your whole history.
Topics with fewer than 3 prompts hold their rates. Slices that thin are too noisy to trust.
Every topic also carries a demand keyword (next section). It models the topic's monthly demand and feeds the Visibility Score's weights.
Topic demand estimates (modeled)
Nobody publishes how often people ask AI assistants anything. So every demand number we show is a modeled estimate, clearly badged as such, and this is the whole model.
Demand is modeled per topic, not per prompt. Real prompts are wordy, and individually they have no measurable search volume.
Every topic carries one demand keyword: the plain head term a person would type into Google for what the topic covers. The topic "Keyword research" carries "keyword research tool".
A language model proposes the keyword at setup, and volumes refresh automatically. A project owner can change a topic's keyword when the automatic pick misses, and the keyword in use is named wherever its number shows.
The math: we take the keyword's Google monthly search volume from the Keywords Everywhere API and divide by 8. That 8 is an empirical average of the ratio between Google search volume and ChatGPT prompt volume for the same topic.
The ratio varies by topic, so treat every estimate as directional: right within a factor of a few, never precise. That is also why we round to two significant figures on purpose.
A project's total demand is the plain sum of its topic estimates.
By country, it works three ways. United States projects use US Google volumes directly.
Projects in the other countries Keywords Everywhere covers (the United Kingdom, Canada, Australia, India, New Zealand and South Africa) use that country's volume when it exists. When it reads zero, we fall back to the global volume scaled by the country's share of its language's online population.
Projects everywhere else use that language-aware fallback directly: the global volume of the localized term, scaled by the country's internet users divided by the internet users of all countries where the answer language is primary. The data behind that is World Bank internet usage plus a curated primary-language mapping.
Why language-aware instead of a share of world population? Because a German term's global volume is already mostly German-speaking demand.
Three honesty rules. A topic has exactly ONE demand keyword, and keywords never repeat within a project, so no demand is ever counted twice.
Keywords with zero Google volume get no estimate at all, rather than a made-up one. And estimates that round below one per month show as unknown.
Volumes refresh monthly.
What ChatGPT searched
ChatGPT often runs its own web searches before answering. The vendor records those queries and everything they retrieved, and we store both with each answer.
An answer with no recorded searches was answered from the model's own knowledge. That is how the dashboard splits "from memory" from "live web search".
Retrieved-but-not-cited pages come straight from that data: pages ChatGPT's search surfaced where you were absent.
Known limitations
We do not model personalization or ChatGPT memory.
Each project collects from one fixed country and language, chosen at creation. The public live-demo pages are collected from the United States.
Consumer answers can differ from API answers. Seven-sample intervals are wide by design.
Sentiment and descriptive wording are language-model classifications, with the floors stated above.
Today we track ChatGPT only.
How we collect
Answer collection runs on a licensed third-party vendor's infrastructure, not our own scrapers.
When we fetch anything directly from your site (your homepage for prompt suggestions, or your robots.txt), we identify ourselves honestly and respect your robots.txt.
We never store your website visitors' data.
If your brand appears in our live demo and you want a correction or removal, request it here.
See where your brand stands in AI
Track how AI talks about your brand, free. 40 prompts, no credit card.
Start tracking free
