Measure AI visibility with a fixed, documented question sample, repeatable collection conditions and answer-level evidence. Report mentions, owned-domain citations and recommendations separately, using completed eligible answers as each visibility denominator. Show planned and unavailable observations beside them. Separate branded from unbranded questions and compare like-for-like segments over time. Repeated observations reveal variability in the sampled conditions; they are not independent people or audience reach. Use the results to investigate specific information gaps, while keeping actual visits, inquiries and sales in separate reporting layers.
Write the specification before the baseline
Choose the audience and buying decisions the sample represents. Record exact question text and assign persistent IDs. Include branded and unbranded questions for distinct purposes rather than mixing them into one unqualified score.
Specify platform, interface, search behavior, language, market and dates. A change in the collection method can change the result; preserve these conditions so the next review can compare like with like.
Define each metric
| Measure | Numerator | Denominator |
|---|---|---|
| Mention rate | Completed answers containing the brand under the naming rule | Completed eligible answers |
| Owned-domain citation rate | Completed answers linking to the agreed domain | Completed eligible answers |
| Recommendation rate | Completed answers satisfying the recommendation rule | Completed eligible answers |
| Collection completion | Completed observations | Planned observations |
Count an answer once per measure, even if it mentions the brand several times. Show the unavailable total separately and explain exclusions. Do not add overlapping rates together.
Preserve the underlying record
Keep the answer, timestamp and source links where available. Record whether the source is owned, independent or a competitor domain. An analyst should be able to inspect why a response was classified as a recommendation or an accuracy issue.
Automatic labels need review when context matters. A negative comparison can mention a brand without recommending it. A citation can support a factual statement without endorsing the company.
Report changes without overclaiming
Use comparable samples and explain changes to prompts, platforms or markets. Report percentage-point changes separately from relative growth, and include sample counts so the reader can judge the scale.
A before-and-after change does not by itself prove that a content update caused the result. Preserve the publication log and consider other changes. Small samples can be useful for investigation without supporting broad market claims.
Connect the report to the work
Each material finding should identify a question, a supporting record, an affected page or source and a next action. Assign an owner and a review date. This makes measurement part of the operating program.
The worked example demonstrates the arithmetic using explicitly fictional data. Use a separate commercial framework for qualified leads, opportunities and revenue.
Protect the question mix before comparing rates
Assign each question a buying stage and a branded or unbranded label. Freeze those categories for the comparison period. Keep newly discovered questions in an exploratory set until a planned baseline revision; otherwise, adding easier branded questions can improve the headline without improving discovery. Record both the fixed sample and the exploratory findings when they serve different decisions.
Hypothetical example: a baseline has 20 completed answers from category questions and 10 from branded questions. The next review adds 20 branded answers. Even if the overall mention rate rises, the two totals describe different question mixes. Compare the original category and branded groups separately, and show the added questions as new coverage. Do not describe the aggregate movement as a like-for-like gain.
Treat repetition as a check on variability
Preserve a record for every planned question, platform and observation occasion. The $990 diagnostic scope is one market, one language, 15 questions, three agreed platforms and two observations: 15 × 3 × 2 = 90 planned observations. Completion counts depend on what can actually be collected and classified. This arithmetic defines workload, not statistical confidence or the number of people reached.
When a result changes between observations, inspect the answer and cited sources before assigning meaning. Report stable presence, intermittent presence and factual errors as different findings. If unavailable observations cluster on one platform or question type, investigate that collection pattern before pooling results. Repeated answers can be related to one another, so do not infer population-level certainty from a large-looking observation total alone.
Common questions
Should unavailable answers count as no mention?
No. A collection failure is not evidence of brand absence. Show it separately.
Is there one universal visibility score?
Different tools and samples use different rules. Inspect the method and evidence before comparing scores.
