AI search · 11 min read
How to measure AI search visibility with a prompt panel you can defend
A transparent method for measuring how often AI assistants mention and cite your business: building a prompt panel, sampling each question more than once, setting a cadence, reading the numbers with their margin of error, and the limits no tool removes.
Want to know where your own site stands? Run a free Growth Scan
The problem
Why AI visibility is a rate, not a rank
A search results page is stable enough that a position means something. An AI answer is generated each time: ask the same question twice and the wording, the businesses named and the sources cited can change. Personalisation, location and the assistant’s settings add more variation.
So the useful question is not “are we in the answer?” but “in what share of answers to these questions are we named, and are we cited?” That is a proportion, estimated from samples, and it comes with a margin of error.
For Google specifically, its documentation says traffic from AI Overviews and AI Mode is included in Search Console’s Performance report under the Web search type, not broken out on its own. Most other assistants give site owners no equivalent report. A prompt panel fills that gap.
Three metrics
Mention, citation and accuracy are different measurements
- Mention rate
- Share of answers that name the business at all.
- Citation rate
- Share that link to or attribute a page on your own site.
- Accuracy
- Share of mentions that describe you correctly: services, prices, places.
The method
Build and run a prompt panel
Every step is written down, so the next period’s number is comparable with this one.
Write the panel from real buyer questions
Use the questions customers ask your sales team, the searches in Search Console, and the comparisons buyers make. Cover the stages: problem questions, category questions (“best X in Y”), comparison questions and brand questions. Do not write questions that name you unless you are measuring brand accuracy.
You get: A numbered prompt list, grouped by stage
Fix the conditions
Decide which assistants, which mode (with or without web search), which location and language, and whether you are logged in. Record them. A change of conditions is a new baseline, not a trend.
You get: A one-page run specification
Run each prompt more than once
Answers vary, so a single run is an anecdote. Repeat each prompt several times per period, in fresh sessions, and record every answer in full rather than a summary.
You get: A log of raw answers with dates
Score each answer the same way
For each answer: named or not, cited or not (and which URL), described accurately or not, and which competitors appear. Write the scoring rules down before scoring, and keep a sample of answers for a second person to re-score.
You get: A scored table, one row per answer
Report rates with their margin of error
Mention rate, citation rate and accuracy per stage and per assistant, each with the number of answers behind it and an approximate margin of error. Flag any change smaller than the margin as noise.
You get: A dated report you can compare period to period
Keep the panel stable, and version it
Add questions when the business changes, but keep a core set unchanged across periods so the trend means something. Record every change to the panel with its date.
The arithmetic
How much a rate can move by chance
Illustrative arithmetic, not observed data: the approximate 95% margin of error for a measured rate of 50%, using the normal approximation for a proportion described in the NIST handbook.
| Dimension | Scored answers | Approximate margin of error | What it means |
|---|---|---|---|
| One question, 10 runs | 10 | About ±31 points | Too wide to read a trend from one question. |
| One question, 40 runs | 40 | About ±15 points | Useful for large changes only. |
| A stage of 20 questions, 5 runs each | 100 | About ±10 points | A reasonable floor for a monthly report. |
| Full panel, 400 answers | 400 | About ±5 points | Small changes start to be readable. |
Limits
What no method or tool removes
Report these alongside the numbers. A measurement that hides its limits invites decisions it cannot support.
- Your panel is not everyone’s questions. Real users ask things you did not think to include
- Real users are personalised, located and logged in differently from your test runs
- Assistants change models and retrieval without notice, which can move every number at once
- Answers from runs are not answers to real users, so rates are not a traffic forecast
- Repeated runs within one session are not independent; use fresh sessions
- A named business is not a recommended one. Read what the answer says about you
Using the numbers
What to do with the results
Start with accuracy. If assistants describe your services, prices or locations wrongly, fix the sources they draw on — your own pages, your Business Profile, the directories and comparison pages that describe you — before chasing a higher mention rate.
Then look at the stages where competitors are named and you are not. Those are content and coverage questions, and the published research on generative engines suggests that content changes can shift how often a source appears — the GEO paper by Aggarwal and colleagues is one widely referenced study — though its test conditions are not your market.
BOOSTD’s platform plans run this kind of sampled check on a schedule, and the pricing page states how often each plan checks and across how many assistants. Whether you use a tool or a spreadsheet, the method above is the test to hold it to.
Questions
Common questions about measuring AI visibility
Can AI search visibility be measured?
Yes, as an estimate. You can measure how often a set of questions produces an answer that mentions or cites you, across repeated runs. You cannot measure it the way a rank tracker measures a position, because the answer varies between runs, users and days.
Does Google Search Console show AI Overviews traffic?
Google’s documentation says clicks and impressions from AI features, including AI Overviews and AI Mode, are included in the Performance report under the Web search type. They are not reported separately, so Search Console tells you the combined total, not the AI share.
How many prompts do I need?
Enough to cover the questions your buyers ask at each stage — often a few dozen for a single-service business — and enough runs of each to make the rate meaningful. The margin of error falls with the square root of the number of runs, so quadrupling runs halves it.
How often should I measure?
Monthly is a reasonable default for most businesses; weekly is useful during a launch or after a significant change. More frequent checks are only worth it if each check has enough runs to be read on its own.
How do I judge a vendor’s AI visibility score?
Ask for the prompt list, the assistants and settings used, the number of runs per prompt, the date range, and how mention, citation and accuracy are each counted. A score without those is not a measurement you can compare over time.
Want a baseline you can compare month to month?
A prompt panel written for your market, run across several assistants on a fixed schedule, and reported with the number of answers behind each rate.
References
Sources
The primary documents and published research this page relies on. Platform rules change, so check the source before acting on a detail.
Related
Where to go next
- how assistants choose sources
- what AI Overviews do to clicks
- AEO, GEO and SEO compared
- ChatGPT and Google search compared
- AI share of voice, defined
- what counts as a citation in an AI answer
- how often each plan checks AI visibilityThe cadence and number of assistants per plan are published.
- how the platform runs the checks
Last updated