AI Share of Voice: A Useful Idea With No Standard Methodology

AI share of voice measures how often your brand appears in AI answers versus competitors. There is no agreed method, vendors measure incompatible things, and the number is not comparable across tools.

Share of voice is a good idea imported from a world that had a denominator.

In paid media it works because impressions are counted by a system both buyers and sellers agreed to. In AI search, nobody agreed to anything. There is no impression log, no published prompt distribution, no query volume data from the engines, and no standard for what counts as an appearance. Every AI share of voice number you have ever seen was produced by a private method somebody chose.

That does not make the metric useless. It makes it an internal instrument rather than a market fact, and the difference matters when someone puts it on a slide.

What it is trying to measure

Run a set of prompts against an AI engine. Count how many answers mention your brand. Compare against competitors. Express as a percentage.

The intent is sound. If eight of your ten category prompts name a competitor and one names you, that is a real signal about how the engine models your category. Directionally it tells you something no rank tracker can.

The problem is every word in that paragraph hides a decision.

The five decisions that change the number

1. The prompt set. This is the single biggest lever and it is entirely arbitrary. Nobody publishes what people actually ask ChatGPT about your category. So the vendor invents a list, or you do. Add five long-tail prompts where you are strong and your share of voice rises. Nothing about your business changed.

2. What counts as an appearance. A linked citation is not the same as a brand name in prose, which is not the same as your domain in a source list the user never expands. Some tools count all three, some count one. Some count a brand mentioned three times in one answer as three, some as one.

3. Which engine and which version. ChatGPT, Perplexity, AI Overviews and AI Mode are different pipelines with different indexes. Answers also move between model versions. A share of voice number without a named model and date is not reproducible even in principle.

4. Sampling. Generative output varies run to run. One sample per prompt per month is noise. Multiple samples per prompt smooths it and costs more. Vendors sit at different points on that trade-off and rarely disclose where.

5. Personalisation, region and history. Answers differ by location and by session context. Measurement environments vary from clean API calls to logged-in browser sessions, and those are not the same product.

Change any one of those and the number moves. There is no adjudicator, because none of the engines publish the data that would settle it.

What the engines actually tell you

Very little, and that is the honest state of it.

Google says sites appearing in AI features are included in overall search traffic in Search Console, inside the Performance report’s “Web” search type. So there is no first-party breakout of AI Overviews impressions to compare a vendor number against. Perplexity and OpenAI publish crawler documentation, not appearance data.

This is the crux. Every AI share of voice product is inferring, from outside, a distribution the engine has not published. That is legitimate research. It is not a measurement standard.

How to use it anyway

You can get real value out of this if you stop treating it as a public number.

Fix your own method and never change it silently. Write down the prompt list, the engine, the model version, the cadence, the region and the counting rule. Version it. If you change the prompt set, break the trend line visibly rather than letting the series absorb the change.

Track the delta, not the level. “We were mentioned in 12% of our 60 prompts in June and 19% in August” is a real finding. “Our AI share of voice is 19%” is a claim about a market nobody has surveyed.

Watch competitor movement more than your own. A rival appearing in prompts they were absent from is a strong signal about something they did. That is the most actionable output of the whole exercise.

Segment by prompt intent. Aggregate share of voice hides the interesting structure. Being absent from “best X for Y” prompts is a different problem from being absent from “how does X work” prompts, and they have different fixes.

Pair it with a cause. Share of voice is a symptom metric. It tells you the score, never why. Diagnosis lives in the crawl and grounding chain, the citations you have earned, and whether the search crawlers can even reach you. Check that last one against the AI crawler directory before you spend a quarter on content.

Be careful with vendor comparisons

If two tools disagree by 20 points on the same brand, that is expected behaviour, not a bug in one of them. Before you pick a tool on the number it produces, ask for the prompt set, the sampling rate, the model version and the counting rule. A vendor that will not answer those four questions is selling you a number you cannot defend.

For a comparison of what the tools in this space actually do, see AI visibility tools compared.

The short version

AI share of voice is a defensible internal trend and an indefensible market statistic. Pick a method, freeze it, track the movement, and never compare your number to one produced by a different tool.

Anyone quoting you an industry-wide AI share of voice figure has made up a denominator. The engines have not published one.

Your check is running.