We measure AI search visibility by running a fixed set of buyer prompts on each AI engine 3 times per cycle, then counting how often each answer names, cites or recommends the brand, compared with the same counts for its competitors.
That sentence hides most of the difficulty. AI answers vary between runs, differ by country and language, change when a model is updated, and are easy to misread from a single screenshot. This article explains how we handle each of those problems, so you can judge our numbers or build the same measurement yourself with a spreadsheet. The formal version of the method is on our measurement page.
What we count
Every answer gets the same five labels. The rules are written down before the first run, and they do not change within a reporting period.
| Label | Rule |
|---|---|
| Mention | The answer text names the brand or an unambiguous variant, such as a product name. |
| Citation | The answer links to or lists the brand's domain as a source, whether or not the text names the brand. |
| Recommendation | The answer presents the brand as an option the user should consider for their case. |
| Position | When the answer lists more than one brand, the brand's place in that list. |
| Error | The answer states something false about the brand: price, product, location, people. |
Two cases are excluded on purpose. A brand named only because the prompt contained it is not a mention; brand prompts are scored separately. A different company with a similar name is not a mention of yours.
The metrics
From the labels we compute a small set of rates. Each one answers a different question.
- Mention rate = answers naming the brand ÷ all answers. Are we in the conversation?
- Citation rate = answers citing the brand's domain ÷ all answers. Is our site treated as a source?
- Recommendation rate = answers recommending the brand ÷ all answers. Are we suggested, or only mentioned?
- Share of voice = the brand's mentions ÷ mentions of the brand plus its tracked competitors. How do we compare?
- Accuracy = mentions without an error ÷ all mentions. Is what they say about us true?
A worked example
This is a hypothetical example, not client data. Suppose a prompt set has 30 prompts, each run 3 times on one engine: 90 answers. The brand is named in 18 of them, so its mention rate on that engine is 18 ÷ 90 = 20%. Across the same 90 answers, the brand and its three tracked competitors are named 120 times in total, 18 of them for the brand. Its share of voice is 18 ÷ 120 = 15%. Two of the 18 mentions state an outdated price, so accuracy is 16 ÷ 18, about 89%.
We report each metric per engine and per market. We do not blend engines into a single score. A brand can be strong in Perplexity and absent from Google AI Overviews, and a blended score would hide exactly the gap you need to see.
Building the prompt set
The prompt set decides what the numbers mean. A set of prompts that already contain the brand name will produce a flattering mention rate that says nothing about discovery. We build ours from the buyer's side.
- Collect real questions from sales calls, support tickets, search console queries phrased as questions, and the communities where buyers talk.
- Group by intent: discovery (“best X for Y”), comparison (“A vs B”), problem (“how do I…”), brand (“is A reliable?”) and local (“X in city”).
- Keep brand prompts separate. They measure what engines say about you, not whether they find you.
- Fix the wording. Small changes in wording can change the answer. Each prompt is stored exactly as asked.
- Split core and exploratory. A core set stays fixed for the reporting period, so trends are comparable. A smaller exploratory set rotates to catch new questions.
- Version it. Every change to the set is logged with a date and a reason.
Why 3 runs per prompt
AI answers are not deterministic. Ask the same question twice and the list of brands can change. A single run can put a brand in or out of an answer by chance. Running each prompt more than once and reporting a rate smooths that out.
We use 3 runs per prompt per engine as the default because it balances cost against stability for most prompt sets. For prompts whose answers swing widely from run to run, we raise the number and say so in the report. If you are measuring on your own, 3 runs is a sensible floor; 1 run is not a measurement.
Keeping runs comparable
A number only means something next to another number collected the same way. We control what we can and record what we cannot.
| Variable | What we do |
|---|---|
| Personal history | Clean sessions with no prior conversation, where the platform allows it |
| Country and language | Fixed per market; each market is measured separately |
| Surface | Consumer app, web and API can answer differently, so we record which one and never mix them in one metric |
| Model | The model name is stored with each answer where the platform shows it |
| Date | Every answer is time-stamped; each cycle runs within a short window |
| Classification | Written rules, applied the same way each time, with a sample checked by a second person |
When a model is updated
Vendors update their models and search systems, not always with public notice. An update can move every number in a report overnight, for reasons that have nothing to do with your work. We handle it in five steps:
- Detect it from the vendor's release notes, or from a sudden shift across the whole core set at once.
- Re-run the full set on the new model.
- Mark a break in the trend line. Numbers before and after are not compared as if nothing changed.
- Re-map the sources the new model cites for your topics.
- Adjust the plan if the sources changed.
What a report shows
- Each metric per engine and per market, with the previous cycle next to it.
- The prompt set version used, and any changes since the last report.
- Every answer behind the numbers, so each count can be checked.
- The domains cited for your topics, and which of them mention you.
- Every factual error found, and its likely source.
- What we changed during the cycle, and what we will work on next.
Limits of the method
This measurement describes a sample of answers under controlled conditions. It does not describe every answer every user sees. Logged-in users with long histories may get different answers. We cannot see inside the models, so when a number moves we separate what the data shows from what we suspect. We say which is which.
Doing it yourself
You can run a simple version with a spreadsheet. Use one row per answer and these columns: date, engine, surface, market, prompt ID, prompt text, run number, brands named (in order), domains cited, mention (yes/no), citation (yes/no), recommendation (yes/no), error (text). Start with 20 to 30 prompts and 3 runs each on the engines your buyers use. Repeat monthly with the same prompts. The rates in this article can be computed with basic spreadsheet formulas.
If you would rather have it done for you, our free GEO audit runs this method on your category.