The short answer: AI share of voice is the percentage of brand mentions your brand earns, versus competitors, when people ask AI engines — ChatGPT, Gemini, Perplexity, Claude — questions in your category. If AI engines mention brands 200 times across a representative set of buyer prompts and 30 of those mentions are yours, your AI share of voice is 15%. It's the AI-era successor to two of the most validated metrics in marketing science: share of voice and share of search.
What is AI share of voice?
AI share of voice (sometimes written "AI SOV" or "share of voice in AI search") measures how much of the conversation your brand owns inside AI-generated answers.
When someone asks ChatGPT "what's the best [your category] for a mid-size UK business?", the answer names a handful of brands. Do the same across Gemini, Perplexity and Claude, across the dozens of question variations real buyers actually ask, and a pattern emerges: some brands dominate the answers, some appear occasionally, and most never appear at all. AI share of voice puts a number on where you sit in that pattern.
We saw this exact pattern when we ran 500 ChatGPT prompts about UK accountants: three software brands appeared in over half of all answers, while the overwhelming majority of accountancy businesses appeared in none.
It's worth distinguishing two metrics that get conflated:
Mention rate is the percentage of prompts where your brand appears at all ("we showed up in 40% of answers"). This tells you about your own visibility in isolation.
AI share of voice is your share of total brand mentions across those prompts. This tells you where you stand relative to competitors — which is what actually determines whether you win the recommendation.
Both matter. But share of voice is the strategic one, because AI answers are a zero-sum shortlist: every answer that recommends three competitors and not you is a buying decision forming without you in the room.
Why share of voice predicts growth (forty years of evidence)
Here's the part that separates AI share of voice from the parade of vanity metrics the industry invents every few years: share of voice is one of the most rigorously validated predictive metrics in the history of marketing.
Les Binet and Peter Field's analysis of the IPA Databank — nearly 1,000 effectiveness case studies across three decades, published in Marketing in the Era of Accountability (2007), The Long and the Short of It (2013) and Effectiveness in Context (2018) — established what's become known as the ESOV rule: for every 10 percentage points of excess share of voice (share of voice minus share of market), a brand gains roughly 0.5% of market share per year.1 Nielsen's independent analysis of 123 brands landed on essentially the same number.2 The LinkedIn B2B Institute later confirmed the relationship holds in B2B, at around 0.6% annual growth per 10 points.3
Brands that command more of the conversation than their market share justifies tend to grow. Brands that command less tend to shrink. It has held across categories, countries and decades.
The metric has already migrated channels once. As search became the front door to buying, the IPA's Share of Search research (led by Binet himself) showed that a brand's share of category search queries works as a cheap, timely leading indicator of market share — the same underlying relationship, measured where attention had moved.4
AI share of voice is the third act of the same story. Attention has moved again — into AI-generated answers, where two thirds of Google searches now end without a click and buyers form shortlists before ever visiting a website. The measurement follows the attention. The logic — own more of the conversation than your size justifies, and growth follows — hasn't changed since Kellogg's out-shouted Post through the Depression.
One honest caveat before anyone over-extrapolates: the 10:0.5 rule was derived from advertising spend data, not AI mentions. Nobody yet has three decades of AI-answer data proving the identical coefficient carries over — the field is two years old. What we do know is that the underlying mechanism (relative salience at the moment of decision drives choice) is precisely what AI answers now mediate, arguably more directly than advertising ever did, because the AI's mention is the shortlist. Treat the rule as strong directional evidence, not a spreadsheet constant.
How to calculate AI share of voice: a worked example
Say you sell accounting software for UK small businesses. You define 30 prompts a real buyer might ask, across three types:
Discovery: "best accounting software for UK small business", "what should a UK sole trader use for bookkeeping"
Comparison: "Xero vs QuickBooks for a small UK company", "alternatives to Sage for startups"
Use-case: "accounting software that handles CIS deductions", "best bookkeeping tool for VAT-registered freelancers"
You run all 30 across ChatGPT and Gemini — 60 answers. Across those answers, brands are mentioned 210 times in total. Your brand accounts for 25 of those mentions.
Your AI share of voice: 25 ÷ 210 = 11.9%.
Run the same count for each competitor and you have the full competitive picture — who owns the category conversation, who's gaining, and which prompt types you're losing. In this example you might find you hold 24% share on use-case prompts (your documentation is strong) but 4% on comparison prompts (you have no comparison content for AI to draw on). That gap is your content strategy, handed to you as arithmetic.
How to get your first baseline (free, in an afternoon)
You don't need a tool for your first measurement — you need a spreadsheet and an honest hour or two:
- Write 20–30 prompts across the three types above, in the language your buyers actually use — your sales team's inbound emails are a goldmine for real phrasing.
- Run every prompt in a fresh session on ChatGPT and Gemini (which between them drive the large majority of AI referral traffic), so personalisation doesn't skew the answers.
- Log every brand mentioned per answer — one row per prompt-engine pair, one column per brand — then compute each brand's mentions ÷ total mentions.
That's a genuine baseline, and for a first look at where you stand it's absolutely worth doing.
Now the honest limitations, because they change what the number means: AI answers are non-deterministic — run the same prompt twice and you'll get variations, so a single pass is a noisy sample, not a measurement. Thirty prompts is a small set. Manual logging drifts as the person doing it changes. And turning a snapshot into what the metric is actually for — trend lines, per-engine breakdowns, repeated sampling for statistical stability, month after month across four engines — stops being an afternoon and becomes a job. That's the genuine (rather than manufactured) reason measurement tools exist for this, the same reason nobody checks rankings by hand any more. Do the baseline manually; automate the moment you need movement over time.
Before building Visibly I ran exactly this exercise manually for brands I worked with — spreadsheet, fresh sessions, prompt list refined from real customer language. Two things surprised me every time. First, how differently the engines behave: a brand dominating ChatGPT answers can be nearly invisible on Gemini, because they build their answer sets from different sources. Second, how fast the numbers move — a competitor publishing one strong comparison page shifted their share within weeks. If you take one thing from the manual method, make it this: measure per-engine, never as a blended average. The blend hides exactly the gaps you need to see.
AI share of voice tools: how the market breaks down
If you'd rather not run spreadsheets, this is an emerging tool category — and it's easier to navigate by tier than by name. Noting the obvious bias that I make one of these, here's the honest map:
Free graders give you an instant single-brand snapshot — HubSpot's AI Search Grader is a good example, and worth running for a first look. What they don't do is competitive tracking or trend lines; they answer "where am I today?" once.
Budget trackers (roughly £50–150/month) monitor your mentions across engines on an ongoing basis. Solid at telling you what your share of voice is; generally light on telling you what to do about it — you get the dashboard, and the interpretation is your job.
Enterprise platforms (£2,000+/month and up) offer the deepest data and integrations, built for Fortune 500 teams with analysts on hand to turn the data into decisions. Overkill in both price and complexity for most in-house teams at mid-size brands.
Visibly — ours — sits deliberately between those last two tiers, built around the view that the number is the start, not the product: share of voice tracked across ChatGPT, Gemini, Perplexity and Claude, competitor comparison, and — the part I built it for — prioritised recommendations for changing the number, not just watching it. See what Visibly does if you want your baseline without a spreadsheet.
The honest buying advice regardless of vendor: if you just want to know where you stand once, use a free grader or the manual method above. Pay for a tool when you need trend lines, competitor tracking and per-engine breakdowns — the things that turn a snapshot into a strategy.
Reading the number: what actually matters
Three nuances that separate useful AI share of voice measurement from a vanity dashboard:
Position within the answer matters. Being the first brand named in a recommendation is not the same as being the seventh in a comparison table. Weight accordingly, or at minimum track "first mention" share separately.
Context matters more than count. A mention can be a recommendation ("X is the best option for…"), a neutral listing, or a caveat ("some users report X is expensive"). Raw share of voice treats all three identically; your revenue does not. As the measurement matures, sentiment-weighted share is where this metric is heading.
Engines genuinely disagree — and the disagreement is the insight. Each engine assembles answers from different sources with different weightings. Your per-engine breakdown tells you which sources to invest in: strong on Perplexity but weak on Gemini usually means your third-party citations are good but your structured, crawlable owned content isn't (or vice versa). A blended average throws that diagnostic away.
What to do with your number
Once you have a baseline, the playbook is the same logic Binet and Field validated, translated to the new channel: find where your share of voice trails what your market position justifies, and close the gap. In practice that means fixing the prompt types you're losing (usually comparison content — the tactics are here), making sure engines can read you at all (free crawler check), and re-measuring on a fixed cycle so you're tracking movement, not snapshots. If you're newer to this whole discipline, the complete GEO guide is the place to start.
The brands winning AI search in five years will be the ones who started measuring when the category was two years old and the incumbents weren't looking. In most categories, right now, nobody owns this conversation yet. That's the opportunity.