Skip to content
BOOSTD

AI search · 11 min read

How to measure AI search visibility with a prompt panel you can defend

A transparent method for measuring how often AI assistants mention and cite your business: building a prompt panel, sampling each question more than once, setting a cadence, reading the numbers with their margin of error, and the limits no tool removes.

Want to know where your own site stands? Run a free Growth Scan

Published by Zubair Afzal (responsible editor), on owner authorisationUpdated Reviewed quarterly — next review:

The problem

Why AI visibility is a rate, not a rank

A search results page is stable enough that a position means something. An AI answer is generated each time: ask the same question twice and the wording, the businesses named and the sources cited can change. Personalisation, location and the assistant’s settings add more variation.

So the useful question is not “are we in the answer?” but “in what share of answers to these questions are we named, and are we cited?” That is a proportion, estimated from samples, and it comes with a margin of error.

For Google specifically, its documentation says traffic from AI Overviews and AI Mode is included in Search Console’s Performance report under the Web search type, not broken out on its own. Most other assistants give site owners no equivalent report. A prompt panel fills that gap.

Three metrics

Mention, citation and accuracy are different measurements

A formula diagram. Useful AI visibility is the product of three terms.
Mention rate
Share of answers that name the business at all.
Citation rate
Share that link to or attribute a page on your own site.
Accuracy
Share of mentions that describe you correctly: services, prices, places.

The method

Build and run a prompt panel

Every step is written down, so the next period’s number is comparable with this one.

  1. Write the panel from real buyer questions

    Use the questions customers ask your sales team, the searches in Search Console, and the comparisons buyers make. Cover the stages: problem questions, category questions (“best X in Y”), comparison questions and brand questions. Do not write questions that name you unless you are measuring brand accuracy.

    You get: A numbered prompt list, grouped by stage

  2. Fix the conditions

    Decide which assistants, which mode (with or without web search), which location and language, and whether you are logged in. Record them. A change of conditions is a new baseline, not a trend.

    You get: A one-page run specification

  3. Run each prompt more than once

    Answers vary, so a single run is an anecdote. Repeat each prompt several times per period, in fresh sessions, and record every answer in full rather than a summary.

    You get: A log of raw answers with dates

  4. Score each answer the same way

    For each answer: named or not, cited or not (and which URL), described accurately or not, and which competitors appear. Write the scoring rules down before scoring, and keep a sample of answers for a second person to re-score.

    You get: A scored table, one row per answer

  5. Report rates with their margin of error

    Mention rate, citation rate and accuracy per stage and per assistant, each with the number of answers behind it and an approximate margin of error. Flag any change smaller than the margin as noise.

    You get: A dated report you can compare period to period

  6. Keep the panel stable, and version it

    Add questions when the business changes, but keep a core set unchanged across periods so the trend means something. Record every change to the panel with its date.

The arithmetic

How much a rate can move by chance

Illustrative arithmetic, not observed data: the approximate 95% margin of error for a measured rate of 50%, using the normal approximation for a proportion described in the NIST handbook.

Approximate 95% margin of error for a 50% mention rate at different numbers of scored answers
DimensionScored answersApproximate margin of errorWhat it means
One question, 10 runs10About ±31 pointsToo wide to read a trend from one question.
One question, 40 runs40About ±15 pointsUseful for large changes only.
A stage of 20 questions, 5 runs each100About ±10 pointsA reasonable floor for a monthly report.
Full panel, 400 answers400About ±5 pointsSmall changes start to be readable.

Limits

What no method or tool removes

Report these alongside the numbers. A measurement that hides its limits invites decisions it cannot support.

  • Your panel is not everyone’s questions. Real users ask things you did not think to include
  • Real users are personalised, located and logged in differently from your test runs
  • Assistants change models and retrieval without notice, which can move every number at once
  • Answers from runs are not answers to real users, so rates are not a traffic forecast
  • Repeated runs within one session are not independent; use fresh sessions
  • A named business is not a recommended one. Read what the answer says about you

Using the numbers

What to do with the results

Start with accuracy. If assistants describe your services, prices or locations wrongly, fix the sources they draw on — your own pages, your Business Profile, the directories and comparison pages that describe you — before chasing a higher mention rate.

Then look at the stages where competitors are named and you are not. Those are content and coverage questions, and the published research on generative engines suggests that content changes can shift how often a source appears — the GEO paper by Aggarwal and colleagues is one widely referenced study — though its test conditions are not your market.

BOOSTD’s platform plans run this kind of sampled check on a schedule, and the pricing page states how often each plan checks and across how many assistants. Whether you use a tool or a spreadsheet, the method above is the test to hold it to.

Questions

Common questions about measuring AI visibility

Can AI search visibility be measured?

Yes, as an estimate. You can measure how often a set of questions produces an answer that mentions or cites you, across repeated runs. You cannot measure it the way a rank tracker measures a position, because the answer varies between runs, users and days.

Does Google Search Console show AI Overviews traffic?

Google’s documentation says clicks and impressions from AI features, including AI Overviews and AI Mode, are included in the Performance report under the Web search type. They are not reported separately, so Search Console tells you the combined total, not the AI share.

How many prompts do I need?

Enough to cover the questions your buyers ask at each stage — often a few dozen for a single-service business — and enough runs of each to make the rate meaningful. The margin of error falls with the square root of the number of runs, so quadrupling runs halves it.

How often should I measure?

Monthly is a reasonable default for most businesses; weekly is useful during a launch or after a significant change. More frequent checks are only worth it if each check has enough runs to be read on its own.

How do I judge a vendor’s AI visibility score?

Ask for the prompt list, the assistants and settings used, the number of runs per prompt, the date range, and how mention, citation and accuracy are each counted. A score without those is not a measurement you can compare over time.

Want a baseline you can compare month to month?

A prompt panel written for your market, run across several assistants on a fixed schedule, and reported with the number of answers behind each rate.

References

Sources

The primary documents and published research this page relies on. Platform rules change, so check the source before acting on a detail.

  1. Google Search Central: AI features and your website
  2. NIST/SEMATECH e-Handbook of Statistical Methods: Confidence intervals for a proportion
  3. Aggarwal et al., GEO: Generative Engine Optimization (arXiv:2311.09735)

Last updated