AI share of voice
AI share of voice is the share of AI answers, across a fixed set of prompts and many repeated runs, in which a brand is named or cited, compared with its competitors.
AI share of voice measures how often AI assistants name or cite your brand, relative to competitors, across a fixed set of category prompts run many times. It borrows its name from advertising share of voice, but it counts appearances in answers, not spend. Because answers change from run to run, the number needs repeated sampling and a stated margin of error.
- Origin
- Practice since 2024, borrowing advertising share of voice (James Schroer, HBR, 1990; Les Binet and Peter Field, IPA), 2024 onward (share of voice: 1990)
- Level
- 401 · Expert
- Fits
- Small and mid-size, Scale-up
- Time to apply
- one day to build the prompt set, then a weekly or monthly run
- What you need
- a list of 3 to 5 direct competitors, with the exact spellings people use for each · 20 to 50 prompts that real buyers would type, grouped by buying situation · access to the assistants you want to track, or a tool that runs them and stores every raw answer
AI share of voice is the share of AI-generated answers, across a fixed set of prompts, in which a brand is named or cited, set against its competitors. It is a measure of presence in answers from assistants such as ChatGPT, Claude, Gemini and Google’s AI Overviews. The label is new and the method is not standardized. What follows is a way to measure it that holds up, built on advertising research that is decades older and on what is known about how language models vary.
Where the name comes from
Classic share of voice (SOV) is a brand’s share of advertising spend in its category. James Schroer’s January 1990 Harvard Business Review article described the share of voice effect: in most consumer markets budgets are stable, and a challenger needs to outspend the biggest rival by at least 100% to gain ground.
The evidence behind the idea comes from the IPA, the UK advertising institute. Les Binet and Peter Field defined extra share of voice (ESOV) as SOV minus share of market (SOM) in Effectiveness in Context. Their model says brands whose SOV sits above SOM tend to grow, and brands below it tend to shrink. In their IPA Databank of campaigns from 1998 to 2016, offline brands gained about 0.5 points of market share per 10 points of ESOV, and online brands about 1.3. The authors note the data are award entries, so the pattern is observational, not experimental. The long and short of it covers that caveat in detail.
Share of search is the nearest digital cousin. Binet presented it in October 2020: a brand’s share of organic Google searches in its category. He reported that it led market share in three categories, by up to a year for cars, and called it “by no means a silver bullet.” AI share of voice has no such track record. Treat it as a visibility measure until someone shows it predicts sales.
What exactly do you count?
Pick one unit and one denominator, then keep them. The cleanest set is three rates, each with its own denominator.
| Measure | Counts | Denominator |
|---|---|---|
| Mention rate | Answers that name the brand | All answers collected |
| Citation rate | Answers that link your domain | All answers collected |
| Share of voice | The brand’s mentions | Mentions of all tracked brands |
Mentions and citations are different events, and a brand can have either without the other.

Citations are the easier one to capture. OpenAI’s web search tool returns url_citation annotations with a URL, title and text position, and Anthropic’s web search tool always returns citations with url, title and a cited-text excerpt. Google’s Generative AI performance report in Search Console, described in Google’s AI features documentation, reports impressions for links shown in AI Overviews and AI Mode. Per Google’s own guide, this is how to see how your content performs there, and third-party tools cannot see inside Google’s systems.
A citation also does not guarantee support. Liu, Zhang and Liang audited four generative search engines in 2023 and found 51.5% of sentences were fully supported by their citations and 74.5% of citations supported their sentence. The Tow Center’s March 2025 test of eight tools on news excerpts found more than 60% of queries answered incorrectly. Log the answer text too, so you can read how the brand is described.
How do you build the prompt set?
Build it from buying situations, not keywords. The Ehrenberg-Bass Institute’s category entry points are the moments and cues that bring a category to mind. Each is a prompt family: a problem, a budget, a use case, a location. Add the wording customers use, taken from sales calls and reviews, because prompts written by marketers drift towards their own vocabulary. See category entry points and voice of customer.
Wording matters more than you would expect. In SparkToro’s study, 142 people given the same headphone need wrote prompts with an average semantic similarity of 0.081. Sclar and colleagues found open models changing accuracy by up to 76 points from format changes that kept meaning. Include several phrasings of every prompt, and tag each prompt as unbranded, comparison or branded. Branded prompts show recall; unbranded ones show discovery.
Why one run tells you almost nothing
The same prompt gives different answers, for three separate reasons.
- The model. Atil and colleagues ran five LLMs set to be deterministic, 10 times each over eight tasks. None gave identical outputs across tasks, and accuracy varied by up to 15% between runs. Thinking Machines traced much of it to server load changing batch size, and in a test of 1,000 completions at temperature 0 got 80 unique outputs.
- The search step. Google describes AI Mode as issuing multiple related searches across subtopics, so the pages feeding an answer can change.
- The assistant’s mix of sources. Chen and colleagues found AI search tools lean towards third-party earned media over brand-owned content, and differ in phrasing sensitivity and freshness.
SparkToro and Gumshoe.ai asked about 600 volunteers to run 12 prompts 2,961 times on ChatGPT, Claude and Google. A repeated brand list was under a 1-in-100 event, and identical order under 1 in 1,000. A brand’s visibility across runs was steadier than its rank. The study is not peer reviewed, its authors say they are not credentialed researchers, and Gumshoe sells tracking, so read it as a strong warning and not as a measurement standard.
The practical answer is repeated sampling and an interval. The Wilson interval, recommended by Brown, Cai and DasGupta over the common Wald formula, suits small samples. A brand named in 6 of 10 runs has a 95% interval of about 31% to 83%. At 100 runs and a true rate of 50%, the range tightens to about 9.6 points either side.

Miller’s error-bar paper makes the same point for model evaluations: report uncertainty, not a bare score. A change counts when intervals separate or a shift repeats over waves.
What vendor tools claim
Tools automate the runs. Ahrefs says Brand Radar tracks mentions and citations across AI Overviews, AI Mode, Perplexity, Copilot, Gemini and ChatGPT, with prompts it says come from real search data and checks daily, weekly or monthly. Semrush says its toolkit reports an AI Visibility Score for how often a brand is mentioned, and lists share of voice in its brand reports without giving a formula. Peec AI says it reports visibility, position and sentiment. Profound says it covers citation analytics, sentiment and a visibility score by platform. These are the vendors’ descriptions and none was independently tested here. Ask each for runs per prompt, raw answer export and an interval, and treat any “rank in AI” as the weak signal SparkToro found it to be.
Do mentions matter if nobody clicks?
Clicks are only part of it. Pew’s analysis of 900 US adults found people clicked a traditional result in 8% of visits with an AI summary against 15% without, and clicked a link in the summary in 1% of visits (Pew Research Center, July 2025). A named brand still shapes who gets considered, which is the logic of mental and physical availability. Getting pages cited is the job of GEO and of content built to be read by machines, as in llms.txt and AI-readable content. Google’s guide states that no special files are required and that its SEO practices still apply.
A Growth Lab plan starts from a baseline like this one, so any content or PR spend can be judged against a measured gap: see Growth Lab.
How to apply AI share of voice, step by step
- Define what counts as a hit. Decide before measuring: a mention is the brand named in the answer text; a citation is a link or source card pointing to your domain. Record both as separate yes or no fields per answer. Result: a written coding rule that two people apply the same way.
- Build the prompt set from buying situations. Write prompts from the moments people start looking: a problem, a comparison, a budget, a country. Take wording from sales calls, support tickets and reviews, and add several phrasings of each. Include unbranded, competitor-comparison and branded prompts, and report them separately. Result: 20 to 50 prompts, each tagged by type.
- Fix the run conditions. Choose assistants, model versions, country, language, logged-out or logged-in state, and whether web search is on. Write them down and keep them constant between waves. Result: a settings sheet attached to every dataset.
- Run each prompt many times and store raw answers. Repeat every prompt at least 30 to 100 times per assistant, spread over several days, because identical runs differ. Save the full answer text and the cited URLs. Result: a table of one row per answer, with no summaries yet.
- Score mention rate, citation rate and share. For each brand, mention rate is answers naming it divided by answers collected. Share of voice is its mentions divided by all tracked-brand mentions. Add a Wilson confidence interval to each rate. Result: a table of rates with ranges, not single scores.
- Compare waves, then act on gaps. Treat a change as real only when the ranges stop overlapping or the shift repeats across waves. Read the sources AI cites where you are absent, and fix the pages and third-party coverage that are missing. Result: a short list of gaps tied to specific prompts.
Examples
A dental clinic chain in one city
Illustrative. The chain tracks 30 prompts such as 'best clinic for dental implants in the city' and 'cost of veneers', each run 50 times on two assistants. Of 3,000 answers (30 prompts, 50 runs, two assistants), the chain is named in 600, a 20% mention rate, and a competitor is named in 1,350, 45%. The chain is cited as a source in only 150 answers. The citation list shows the competitor's pricing page and two local directories, so the team fixes its own price page and its directory profiles first.
A B2B payments provider entering a new market
Illustrative. A payments provider tracks 40 unbranded prompts about cross-border payouts, split into 'what is' questions and 'which provider' questions. It is named in 8% of 'which provider' answers but 30% of 'what is' answers. That gap says buyers see the brand while learning but not while choosing. The team adds comparison pages and seeks third-party reviews, then re-runs the same prompts four weeks later.
When to use it
Use it when buyers research with AI assistants before they contact you, when competitors are cited in answers you want to appear in, or when you need a baseline before investing in content, reviews or digital PR. It suits categories where people ask for recommendations, such as services, software and clinics.
When not to use it
Skip it when buyers rarely use assistants for your category, when you cannot afford repeated runs, or when you would report a single screenshot as a score. Do not treat it as a ranking, and do not use it to predict market share: no published evidence links AI share of voice to sales.
Common mistakes
- Reporting one run as the result. Answers vary from run to run, so a single answer, or a ranking position taken from it, is mostly noise.
- Counting mentions and citations as one number. A brand can be named without a link, and a link can appear without the brand named in the text.
- Testing only branded prompts. A brand that appears when its name is typed shows little about whether buyers who do not know it will meet it.
- Changing settings between waves, such as model, country or web search, then reading the difference as progress.
- Averaging across assistants into a single score. Each assistant draws on different sources, so keep them apart.
FAQ
What is AI share of voice?
It is the share of AI answers in which your brand is named or cited, compared with competitors, across a fixed prompt set run many times. The idea comes from advertising share of voice. The measured quantity differs: it counts appearances in generated answers rather than shares of ad spend.
How many prompts and runs do I need?
No standard exists. SparkToro, whose study is not peer reviewed, advises usually at least 60 to 100 runs per prompt. With 100 answers and a true rate near 50%, a 95% Wilson interval is about 9.6 points either side, so small changes between waves are not reliable.
What is the difference between a mention and a citation?
A mention is the brand named in the answer text. A citation is a link or source card to a page. They can occur separately. Count both, because a mention shapes what a buyer remembers and a citation shows which pages the assistant used.
Is AI share of voice the same as classic share of voice?
No. Classic share of voice is a brand's share of category advertising spend. Binet and Field's IPA data link share of voice above market share to growth. No equivalent evidence exists for AI answers, so treat it as a visibility measure, not a growth predictor.
Can I track AI visibility with a tool?
Yes. Vendors such as Ahrefs, Semrush, Profound and Peec run prompts and report visibility. Their claims about coverage and accuracy are their own. Check how many runs per prompt they use, whether raw answers can be exported, and whether they show intervals.
Sources
- IPA, Presentation: The long and short of it, Les Binet and Peter Field, 2013
- Les Binet and Peter Field, Effectiveness in Context: a manual for brand building, IPA
- James C. Schroer, Ad Spending: Growing Market Share, Harvard Business Review, January 1990
- IPA, Binet presents fast, cheap, predictive Share of Search metric, October 2020
- Pranjal Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024
- Berk Atil et al., Non-Determinism of Deterministic LLM Settings, 2024
- Evan Miller, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations, 2024
- Melanie Sclar et al., Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design, ICLR 2024
- Nelson F. Liu, Tianyi Zhang, Percy Liang, Evaluating Verifiability in Generative Search Engines, Findings of EMNLP 2023
- Horace He and Thinking Machines Lab, Defeating Nondeterminism in LLM Inference, September 2025
- Rand Fishkin, SparkToro, New research: AIs are highly inconsistent when recommending brands or products, January 2026
- Search Engine Journal, AI recommendations change with nearly every query, SparkToro finds, January 2026
- Google Search Central, AI features and your website
- Google Search Central, Guide to optimizing for generative AI features on Google Search
- Google, AI Mode in Google Search, March 2025
- Klaudia Jaźwińska and Aisvarya Chandrasekar, Tow Center for Digital Journalism, AI search has a citation problem, Columbia Journalism Review, March 2025
- Pew Research Center, Google users are less likely to click on links when an AI summary appears, July 2025
- OpenAI, Web search tool, API documentation
- Anthropic, Web search tool, API documentation
- Mahe Chen, Xiaoxuan Wang, Kaiwen Chen, Nick Koudas, Generative Engine Optimization: How to Dominate AI Search, 2025
- Lawrence D. Brown, T. Tony Cai, Anirban DasGupta, Interval Estimation for a Binomial Proportion, Statistical Science 16(2), 2001
- Ehrenberg-Bass Institute, Category entry points dissected: how they really contribute to growth
- Ahrefs, Brand Radar (vendor page)
- Semrush, AI Visibility Toolkit (vendor documentation)
- Profound (vendor page)
- Peec AI (vendor page)
Last updated Oct 9, 2026


