Why the old dashboards can't see this
Every established marketing metric assumes a observable surface: rankings you can crawl, impressions a platform reports, mentions a listener can index. Generative answers have none of that. There is no ranking to crawl: answers are composed per conversation. There is no impression report: the platforms do not tell you when your brand appeared. The surface is only observable by asking, systematically and repeatedly.
That is why AI share of voice has to be built as a sampled metric, the way pollsters build approval ratings. You define a population of prompts, sample answers across platforms on a fixed cadence, and compute your share of the recommendations that come back. Done casually, this produces noise. Done with discipline, it produces a metric stable enough to trend, target, and tie to pipeline.
A working definition
We define AI share of voice as: of all brand recommendations returned across a fixed prompt panel and platform set in a period, the percentage that are yours. Weighted variants adjust for position (a first recommendation counts more than a fifth) and for prompt value, so a high-intent comparison prompt influences the score more than a broad category question.
The definition's power is in what it excludes. It does not count raw mentions, because 'unlike <your brand>, X is affordable' is a mention working against you. It counts recommendations: moments where an engine put you on a buyer's shortlist. In Competitor Watch, that is the number sitting next to each competitor's logo, and it is the one that behaves most like market share.
The three design choices that make it durable
First, fix the prompt panel. Share of voice is only comparable over time if the denominator holds still. Version your panel like code: additions and removals happen deliberately, with a change log, never silently. Second, fix the platform set and weight it honestly: if your buyers live in ChatGPT and Perplexity, a Grok-driven bump should not mask a ChatGPT decline. Blended scores need visible per-platform breakdowns.
Third, fix the cadence and sample repeatedly. Generative answers vary run to run; single samples are anecdotes. Weekly panels with multiple samples per prompt-platform pair (the way Visibility Radar runs its scans) smooth run-level variance into a trendable line, with a confidence band that tells you when a movement exceeds normal noise.
Surviving model updates
The hardest measurement problem in this space is that the instrument keeps changing. Platforms ship new model versions, retrieval changes, answer formats shift, and share-of-voice lines jump for reasons that have nothing to do with your marketing. The framework has to expect this rather than be embarrassed by it.
Two practices help. Annotate the timeline: when a platform ships a known model update, mark it on the chart, and evaluate your trend within regimes rather than across them. And watch relative position, not just absolute score: if a model update drops everyone's citation rate but your share of the remaining recommendations holds, your competitive position is intact even though the raw line dipped. Share of voice is robust to instrument change precisely because it is a ratio; absolute visibility scores are not.
From measurement to movement
A durable metric earns a seat in the executive dashboard, and that is where this one belongs: next to pipeline and brand search volume, reported from the Command Center with the same cadence and confidence as any revenue metric. The teams doing this well set quarterly share-of-voice targets per platform, the way they once set ranking targets.
The metric also earns its keep diagnostically. When share of voice moves, the underlying scan data says why: which prompts flipped, which citations changed, which competitor gained. That drill-down path (from the number your CMO watches to the specific source an engine started citing last Tuesday) is what separates a measurement program from a vanity chart.
Published May 12, 2026