Methodology
How we collect the data, how we count a mention, how often we ask, and what we cannot measure. Everything on this page is generated from the code that runs in production, so it cannot quietly go out of date.
We publish this because the strongest criticism of AI-visibility tools is a fair one: the numbers move, and most vendors will not tell you how theirs are made. You should be able to check ours.
How we collect
We send your prompts to each engine’s official API and read the answer it returns, with its citations. We do not scrape, and we do not simulate a browser. Every answer is stored exactly as received, forever, and nothing is ever overwritten — every metric on your dashboard can be recomputed from the raw answers at any time.
Each engine carries a confidence tier, shown next to every number in the product. We never blend them into one opaque score:
Measured — Perplexity
Measured directly from the engine's API with native citations (ground truth).
Estimated — ChatGPT, Claude, Gemini
Estimated via an API proxy with web search. This is NOT what a user sees in the app: the API uses a different system prompt and retrieval layer, so the brands and sources it returns differ substantially. It is a consistent, repeatable proxy — not a replica.
Directional — Google AI Overviews
Directional signal from SERP data; the AI Overview may not trigger on every query.
Exactly which model answers
Not “ChatGPT” — the specific model, named. Answers are stamped with the model version that produced them, so a measurement can never be silently attributed to a model that did not make it.
| Engine | Model |
|---|---|
| Perplexity | sonar-pro |
| ChatGPT | gpt-4.1 |
| Claude | claude-opus-4-8 |
| Gemini | gemini-3.5-flash |
| Google AI Overviews | dataforseo-google-aio |
How often we ask, and how many times
A visibility percentage is only as good as the number of answers behind it. We show that number — n — next to every figure, and a 95% confidence interval with it.
| Plan | Frequency | Answers per prompt, per engine, per run |
|---|---|---|
| Starter | weekly | 3 |
| Pro | daily | 1 |
| Advanced | daily | 1 |
A brand’s first run takes 3 answers per prompt per engine, so your first dashboard is not built on a single reading.
How we decide you were mentioned
We match your brand name and every alias you configure against the text of the answer. A mention is a whole-answer fact: named or not named. We do not infer that you were “probably meant”.
Where an answer puts brands in a numbered or bulleted list, we also record your position in it. We average that position only over answers that actually ranked you — a brand named in prose has no ordinal, and treating it as “last” would poison the average.
Sentiment is scored per brand, per answer — not per answer overall. An answer can name you and still bury you (“X is fine, but Y is faster”), and only per-brand scoring sees that. When we could not score it, we show a dash. We never fill in a neutral 50.
Where we ask from
Every answer freezes the locale and country it was collected under, along with the exact prompt text sent at that moment. Editing a prompt later changes what we ask from then on — it never rewrites what we asked before.
When we change something
Every run, every failure, every configuration change is recorded as a dated event tied to your workspace — including the changes we make. Anything that would alter your numbers (a model change, a frequency change) is visible as a change, not as an unexplained shift in a line chart.
Your raw data
You can export every underlying answer, mention and citation as CSV, and read your metrics from a JSON API with your own key. We do not hold your data hostage to check our work — if our numbers are wrong, you should be able to prove it.
What we cannot measure, and will not pretend to
This section is the point of the page.
AI answers are not deterministic. Ask the same engine the same question twice, minutes apart, and you can get a different list of brands and a substantially different set of sources. This is not our tool being unstable, and it is not something any vendor can fix — it comes from how these models are served. It is exactly why we report a confidence interval instead of a single confident number, and why we will never sell you a “rank position” in an AI answer. A rank implies a stability that does not exist.
An API is not the app. We measure through official APIs. That is not what a logged-in person sees in ChatGPT: the app uses a different system prompt and retrieval layer, and published comparisons find the brands and sources returned differ substantially. Nobody can measure what a logged-in user privately sees — not us, not anyone scraping either. What we offer is a consistent, repeatable proxy, labelled as one.
We do not know how many people ask your prompts. No engine publishes prompt volume. Tools that show you one are modelling it. We would rather show you nothing than a number we invented.
Visibility is not revenue. We measure how often the engines name you and how they describe you. We do not claim to measure what that is worth. No tool in this category has demonstrated that link, and we are not going to be the first to assert it without evidence.
Not measured is never shown as zero. If an engine returned nothing, you see a dash. A zero would claim we asked and nobody named you — a statement about your brand. A dash says we have not measured this — a statement about us. Engines we did not measure are excluded from your overall score rather than dragging it down.
Found something here you think is wrong?
Export your raw answers and check us. That is what the export is for.