What Is AI Share of Voice (Share of Model), and How Do You Measure It?
By Emmanuel Chaumeau — Executive specializing in revenue, transformation and AI-driven value creation
AI share of voice, often called Share of Model, is a comparative measure of the space a brand occupies in responses from assistants such as ChatGPT, Gemini, Claude or Perplexity. It does not measure absolute visibility on the web: across a stable set of traveler scenarios, it measures how often your brand appears, is actually recommended or is cited compared with its competitors. In other words, the right question is not just “is my hotel mentioned?” but “for the same requests, against the same competitors, in the same engine, how often am I present, chosen or supported by a citation?”
AI Share of Voice Is Not Just a Mention Count
Counting how often a brand appears in an AI response is useful, but insufficient. A brand may be mentioned in passing, listed among others or, conversely, clearly recommended as the best choice. These situations do not carry the same value. To make the measurement actionable, you need to distinguish at least three levels: presence, recommendation and citation. Presence answers the question “does the brand appear?” Recommendation answers “is the brand presented as a relevant or preferable option?” Citation answers “does the response explicitly draw on a visible source linked to this brand or its information ecosystem?” If you combine these three levels into a single figure without a clear method, you get an appealing but largely unhelpful indicator. If you separate them, you can understand whether your problem stems from a lack of visibility, a lack of preference or a lack of visible evidence. To clarify these levels of presence in generative responses, you can also read the difference between being cited, visible and recommended by AI.
What AI Share of Voice Actually Measures
AI share of voice measures a competitive position within a defined scope. This scope includes a question corpus, a competitor panel, one or more engines, a language, a market and a testing period. This is essential, because there is no universal AI share of voice. You may have a strong share of voice for requests such as “family hotel near central Rome” and a weak share of voice for “design boutique hotel for a romantic weekend.” You may also be highly visible in English and much less so in French. The metric therefore always depends on the testing scope. It should be read as an indicator of competitive coverage across comparable cases, not as an overall truth about your brand awareness.
Start with a Stable Corpus of Traveler Scenarios
Measurement quality depends first and foremost on the quality of the test corpus. You need to build a set of traveler scenarios stable enough to be rerun over time. A good corpus is not limited to generic queries such as “best hotels in Paris.” It covers concrete scenarios: destination, budget, duration, traveler profile, neighborhood, use case, constraints, seasonality or intent. For example, one request might concern a business trip, another a weekend as a couple, and another a family hotel near a point of interest. The aim is not to test every possible wording, but to test a consistent sample of the requests that really matter to your market. If the corpus changes with every measurement, you are no longer tracking your progress: you are changing the exam each time. To understand why the choice of wording strongly influences results, a useful companion to this topic is why recommendations change depending on the question asked.
Compare Competitors Only on the Same Prompts
A comparison is meaningful only if all brands are observed on exactly the same playing field. You therefore need to test your brand and its competitors on the same prompts, under the same conditions. Comparing your presence across a series of “luxury” questions with a competitor’s presence across a series of “family” questions produces no meaningful signal. Similarly, combining multiple languages, markets and engines into a single average often conceals the real differences. Best practice is to keep results separate, at least initially, by engine, language, market and category of traveler scenarios. You can aggregate them later, but only after understanding what you are aggregating.
Define Coding Rules Before Starting Measurement
Before you even begin testing, you need to decide how you will read the responses. Without a coding framework, two people may interpret the same response differently. A simple framework might include the following categories:
- Absent: the brand does not appear.
- Mentioned: the brand is named without really being proposed as a choice.
- Listed: the brand is included in a selection of options.
- Recommended: the brand is presented as particularly suitable, a priority choice or the best fit for the stated need.
- Cited as a source or supported by a visible source: when the interface displays explicit sources and one of them links to the brand’s website or an external source documenting it.
- Ambiguous: the wording does not allow a clear determination. The important point is not to invent a complex taxonomy, but to apply the same convention consistently. An imperfect but stable metric is better than a theoretically perfect metric coded differently with every run.
The Three Priority Indicators to Track
1. Presence Rate
Presence rate measures how often your brand appears in responses across the tested corpus. Simple formula: presence rate = number of prompts where the brand appears / total number of prompts tested This indicator tells you whether you enter the models’ response space. It does not yet tell you whether you are preferred.
2. Recommendation Rate
Recommendation rate measures how often your brand is actually proposed as a relevant option, rather than merely mentioned. Simple formula: recommendation rate = number of prompts where the brand is recommended / total number of prompts tested This is often the most useful indicator for assessing competitive performance, because it better reflects the shift from simple visibility to choice.
3. Citation Rate
Citation rate can only be measured properly when the interface displays visible sources or references. You then need to decide your calculation convention in advance. There are two options:
- measure citation across all prompts tested;
- measure citation only across prompts where the engine displays sources. The key is to state the denominator clearly. Otherwise, comparisons become misleading, particularly between engines that do not all expose their sources in the same way.
How to Calculate a True Competitive Share of Voice
The three rates above already provide a useful view. To discuss share of voice in the strict sense, you can then convert these observations into a relative share against the competitor panel. The principle is simple: across the same corpus, count how often each brand is present, recommended or cited, then calculate each brand’s share of the observed total. For example, for presence: presence share of voice = appearances of your brand / total appearances of all brands in the panel The same reasoning can be applied to recommendation and citation. This approach has a major advantage: it shows not only your own frequency, but also your relative weight in the competitive space observed. A brand may have a reasonable presence rate yet remain far behind two dominant competitors. Conversely, its relative share may increase even if its raw volume remains stable, because the rest of the panel is declining.
Add Position Within the Response Without Confusing It with Share of Voice
In many responses, presentation order matters. Being mentioned first does not have the same impact as being placed fourth at the end of a list. You can therefore supplement the measurement with indicators such as:
- Top 1 frequency;
- Top 3 frequency;
- average rank when the response is ordered;
- share of cases where a specific competitor appears ahead of you. These indicators are valuable, but they should not replace presence and recommendation rates. First, because not all responses are explicitly ranked. Second, because a conversational response may strongly recommend a brand without presenting it in a numbered list.
Repeat Tests to Reduce the Effect of Variability
An AI response is not a perfectly fixed result. It may vary depending on the engine, language, market, time of testing, session history or exact wording of the request. That is why a rigorous measurement cannot rely on a single attempt. Tests need to be repeated to see whether trends hold. The smaller the gap between two competitors, the more caution is needed. In practice, this means that an isolated gap should not be interpreted as a strategic verdict. What matters is the recurrence of the signal across a consistent corpus and multiple comparable runs. To understand why engines do not behave in the same way, you can read ChatGPT, Gemini, Claude, Perplexity: what are the differences for tourism businesses?.
Separate Results by Engine, Language and Market
An overall average score is tempting, but it often hides the most useful information. A brand may be strong in one engine and weak in another. It may dominate its domestic market and disappear in an international market. It may also be accurately captured in one language and poorly understood in another. The right sequence is therefore:
- measure separately by engine;
- measure separately by language;
- measure separately by market;
- compare categories of traveler scenarios;
- only then produce an overview. This avoids confusing a content problem, a local positioning problem, a translation problem or a simple engine effect.
How to Interpret Gaps
AI share of voice becomes useful when it enables a diagnosis. Low presence often suggests that the brand rarely enters the models’ response space for the scenarios tested. This may indicate a lack of clarity, topical coverage or external corroboration. Reasonable presence but a low recommendation rate suggests something else: the brand is known, but is not considered sufficiently relevant or preferred for the stated need. The problem is no longer simply appearing in responses, but being the right choice. Strong recommendation in only a few areas often reveals specialization. This is not necessarily a bad thing. It may show that the brand has a genuine strength in certain use cases, but not yet across its entire target market. Finally, low citation despite reasonable presence may indicate that the brand appears in responses, but without visible evidence or sufficiently identifiable support in the sources displayed. When measurement shows that a competitor regularly outperforms you, the logical next step is to analyze why AI recommends a competitor rather than your business.
Limitations to State in Every Report
AI share of voice is useful, but it must be presented alongside its limitations. First limitation: it depends on the corpus tested. If you change the traveler scenarios, you also change the metric. Second limitation: it depends on the engine. Two assistants may produce different selections without one being “wrong” and the other “right.” Third limitation: it depends on the response format. Some interfaces cite sources; others do not. Some rank options; others write in natural language. Fourth limitation: it does not directly measure business performance. An increase in AI share of voice is not, on its own, evidence of increased traffic, demand or bookings. Fifth limitation: it remains sensitive to variability. The smaller the gaps, the more important it is to avoid definitive conclusions. Finally, AI share of voice should not be confused with traditional SEO share of voice. The two approaches may overlap, but they do not describe the same thing. A page that ranks well on Google may still have little visibility in a generative response, as explained in why content that ranks well on Google can remain invisible in AI responses.
The Minimum Dashboard to Set Up
Simple tracking is often enough to get started. A minimum dashboard can include:
- presence rate by engine;
- recommendation rate by engine;
- citation rate where the format allows;
- relative share against key competitors;
- a breakdown by language;
- a breakdown by market;
- a breakdown by category of traveler scenarios;
- an indication of stable gaps and gaps that remain volatile. The aim is not to produce an impressive dashboard, but one that supports decisions. If you cannot explain what a change means, the indicator is still too raw.
Key Takeaways
AI share of voice is a comparative metric, not an absolute one. It measures a brand’s place in generative responses against competitors, across a stable corpus of traveler scenarios. To measure it correctly, you need to separate, at a minimum, presence, recommendation and, where possible, citation. You need to test the same prompts for all players, maintain a consistent scope, segment results by engine, language and market, then repeat observations to limit the effect of variability. Used well, this measurement does more than tell you whether you are visible. It shows where you are chosen, where you are not, and in which competitive areas your positioning truly holds up. If you want to move from intuition to a comparable measurement, the useful next step is to set up a comparative GEO audit across a stable corpus of traveler scenarios, to assess presence, recommendation and citation against competitors, engine by engine, language by language and market by market.
