See how often AI engines mention your brand vs. rivals. Arjun Karnik's 7-step AI SOV framework covers ChatGPT, Perplexity & more. Start measuring now.
Written by: Arjun Karnik, Growth Marketing Specialist
Key Takeaways for Measuring AI Share of Voice
AI share of voice (AI SOV) tracks how often AI answers mention or cite your brand versus competitors, replacing rankings as the core visibility metric.
Buyers now rely on AI assistants for vendor decisions, with 69% changing their choice based on AI responses and 33% buying from brands they had not heard of before.
A seven-step playbook covers baseline audits, fan-out query mapping, multi-engine sampling, mention-rate vs competitive SOV calculations, position-weighted scoring, sentiment tracking, and impression-decay tripwires.
Regular measurement and content refreshes are essential, because AI citations decay quickly and pages can lose most visibility within two months without updates.
AI share of voice (AI SOV) is the percentage of AI-generated responses, across a defined prompt set and competitive field, that mention, cite, or recommend your brand relative to all brand mentions in those same responses. The core formula is:
AI SOV = (Your brand mentions ÷ Total brand mentions across all tracked brands) × 100
A brand mentioned in 25 out of 100 answers, while competitors collect 40, 20, and 15 mentions respectively, holds 25% AI SOV. That number only becomes useful when paired with a clear denominator, a documented competitor set, and a locked prompt library. Most teams lack that structure, which makes their AI visibility data noisy and hard to compare.
In under a year the starting point for B2B software research crossed over. More buyers now begin with a chatbot than with Google.
Step 1: Baseline Visibility Audit Across AI Engines
Start by capturing the current state across ChatGPT, Google AI Overviews, Perplexity, and Gemini. The audit shows where your brand appears, where competitors appear instead, and where no brand appears at all.
Run a seed set of 20–30 category prompts across each engine, logging every brand mention, every citation (domain sourced directly), and every instance of competitor presence. This combined dataset becomes your control group, the baseline against which you compare every later run to detect changes in visibility.
Log results in a spreadsheet with columns for engine, prompt, brand mentioned, mention position, citation (yes or no), and sentiment. That structure feeds every downstream step in the playbook.
Step 2: Fan-Out Query Mapping From Buyer Prompts
Each buyer prompt triggers many hidden retrieval queries rather than a single lookup. The AI system assembles the final answer from those underlying queries, so content needs to match the fan-out layer, not just the visible prompt.
Extract fan-out queries directly from ChatGPT instead of inferring them from keyword tools. The target is the machine's questions, not the human's. Use this prompt structure:
"If a B2B buyer asked you '[your category question]', what sub-questions would you need to answer to give a complete response? List them."
Run that extraction across 10–15 seed prompts. The output becomes your production queue and determines which pages you create or rewrite, along with the exact language you use. In Arjun's tests on his own site, pages rewritten to match extracted fan-out queries earned citations while control pages did not.
Step 3: Multi-Engine Sampling Design for Stable Data
Once you have fan-out queries and a structured prompt set, test how your content performs against those prompts across all major AI platforms. Run the full prompt set across ChatGPT, Google AI Overviews, Perplexity, and Gemini.
Lock four variables before any run: the prompt set, the competitor list (five to seven names), the engines, and the weekly cadence. These variables define your measurement instrument, so changing any one of them mid-cycle means you are no longer measuring the same thing and week-over-week comparability breaks.
Step 4: Mention-Rate vs Competitive-SOV Calculation
Mention rate and AI SOV measure different things, and treating them as interchangeable leads to bad decisions.
Metric
Formula
What it measures
Denominator
Mention Rate
Answers mentioning brand ÷ Total valid answers × 100
Absolute appearance frequency
Total answers in the prompt set
AI Share of Voice
Your brand mentions ÷ Total brand mentions across all tracked brands × 100
Mention rate over time against a competitor benchmark. This is the number that replaces rank position. The figures shown are a product view, not a client result.
Step 5: Position-Weighted Scoring Table for Influence
Not all mentions carry the same weight. A brand named first in an AI answer influences buyers more than one listed fourth, so position-weighted SOV applies harmonic decay to reflect that reality.
The standard harmonic decay model assigns weights using the formula Weight = 1 ÷ Position:
Step 7: Impression-Decay Tripwires and a Self-Healing Loop
Once you have baseline metrics for presence, position, and sentiment, the next challenge is keeping those gains over time. Content does not hold its position in AI answers without maintenance, and visibility can erode quickly.
In Arjun's tests on his own site, pages lost the majority of their AI visibility within two months when left unmaintained, which created the decay pattern behind the tripwire thresholds described below. That decay stays invisible unless you instrument for it, and by the time it appears in a monthly report the position is usually gone.
The publishing cadence, running. New articles and refreshes sit in one queue, and pages that have started to slide get flagged and rewritten without anyone auditing a spreadsheet.
Three quarters of cited pages were updated inside a year, and the consistently cited ones averaged under six months. The page you refresh beats the page you write.
Set tripwires in Google Search Console against your baseline impression curve, configured to flag any page that drops more than 20% week-over-week for two consecutive weeks. This threshold is calibrated against the decay behavior Arjun measured on his own site and catches problems early enough that a minor refresh usually fixes them. A refresh does not need to be a full rewrite, because adding a current statistic, updating a date reference, or expanding an answer section is often enough to re-enter the freshness window. 76.4% of pages cited by ChatGPT had been updated within the prior 30 days. The refresh loop becomes a moat, because it is hard for competitors to sustain and easy for incumbents to neglect.
Prompt-Library Construction for Reliable Measurement
The prompt library acts as the measurement instrument, and its quality determines the reliability of every SOV figure downstream. Build it in three layers:
Awareness prompts are category-level questions a buyer asks before they know vendor names. Example: "What tools help B2B companies track AI search visibility?"
Consideration prompts are comparison and evaluation questions. Example: "Compare [your category] options for a $5M SaaS company with two marketers."
Decision prompts are high-intent, specific queries. Example: "Which [category] platform is best for measuring share of voice in AI answers?"
For each awareness prompt in your seed set, you need to extract the underlying fan-out queries that AI systems actually use to construct answers. To extract fan-out queries from ChatGPT directly, use this prompt template:
"You are a B2B buyer researching [category]. List every sub-question you would need answered before recommending a vendor. Format as a numbered list."
Benchmarks vary by category concentration and how many competitors the AI surfaces. As a working framework, below 15% indicates a significant citation gap in most B2B categories, 15% to 25% is a competitive range, above 25% signals strong visibility, and above 40% suggests category leadership. Category leaders rarely exceed 60% because AI systems naturally diversify their citation sources. For B2B SaaS specifically, the top quartile clears 35% or above. The more useful target is relative performance: your AI SOV compared to your current market share. Brands holding higher AI SOV than their market share tend to grow, while those with lower AI SOV tend to shrink, so track the gap between the two numbers rather than the absolute percentage alone.
How often should I re-measure AI SOV?
Plan for weekly spot-checks on 10–15 high-priority prompts and a full prompt-library run monthly. The weekly cadence catches decay before it compounds, while the monthly full run produces the trend data needed for strategic decisions. AI citations are volatile by nature, and the median brand's week-over-week citation count swings by more than 50%, so monthly-only measurement misses most of the signal.
Set impression-decay tripwires in Google Search Console to flag individual pages between scheduled measurement runs. The tripwires catch page-level decay, the weekly spot-checks catch engine-level shifts, and the monthly full run tracks competitive position over time. All three layers matter because they measure different parts of the system.
How do I handle direct and branded-search attribution for AI-driven traffic?
AI-influenced demand often shows up in analytics as direct traffic or branded search rather than anything traceable to the answer that caused it. The buyer reads the AI answer, closes the tab, and types your name directly into a browser or Google, which produces no referral URL and no UTM parameter. Whatever you measure in standard attribution is a floor, not a ceiling.
Three signals help confirm upstream AI influence: unexplained increases in direct visits from new users, rising branded search queries in Google Search Console that do not correlate with active brand campaigns, and traffic arriving via chatgpt.com or equivalent AI referrer domains. Segment AI referrers as a distinct traffic class in analytics, because they convert differently from cold organic traffic and usually arrive pre-educated. Track all three signals together rather than relying on any single attribution model, and add self-reported source fields on high-intent forms as a fourth data point that no analytics tool can capture automatically.
Buyers read an answer, then type your name into a browser. That visit lands as direct or branded search, so whatever you measure here is a floor and never a ceiling.
What is the difference between a citation and a mention in AI answers?
A mention means your brand name appears somewhere in the AI response text, while a citation means your domain was sourced directly as part of the answer, typically with a link or explicit attribution. The two metrics can diverge sharply and measure different things. An AI answer may name your brand without linking to your site, which is common in ChatGPT responses that synthesize from training data. An answer may also cite your URL without naming your brand in prose, which is common in Perplexity, where external links appear in over 77% of responses.
Track both metrics separately. Mention rate tells you how often you are part of the conversation, and citation rate tells you how often your content acts as the evidence behind the answer. Citation rate is the stronger signal for content strategy because it confirms that the retrieval layer is reading and using your pages, not just that your brand name exists in the model's training data.
Conclusion: Run the Loop and Own the Answer Layer
The seven steps above form a closed loop rather than a one-time audit. The baseline audit establishes the starting line. Fan-out query mapping builds the production queue. Multi-engine sampling produces stable data. The two-denominator calculation separates absolute visibility from competitive position. Position-weighted scoring reveals whether mentions are leading or trailing. Sentiment and citation tracking separate presence from influence. Impression-decay tripwires keep the whole system self-healing instead of silently decaying.
In Arjun's tests on his own site, the GEO subfolder went from zero to the only source of new impressions on the entire domain in 60 days, powered by AI Growth Agent at five to eight autonomous actions per day. The fan-out query method described in Step 2 proved decisive, and pages optimized this way consistently earned citations in Arjun's testing. The automated refresh loop described in Step 7 produced the GEO subfolder results mentioned earlier and showed how measurement and maintenance work together as a system.
OpenAI reported 900 million weekly active ChatGPT users in February 2026. Google AI Overviews serves over 2.5 billion monthly active users. The answer layer now shapes buyer shortlists long before sales conversations begin, and it is where your buyers already make early decisions. The window for outsized gains is open now for the same reason it was open in the early SEO era: businesses that decode the new answer layer first lock in the settled answers that later competitors must displace.
The measurement framework in this playbook is fully replicable without proprietary tooling. Run it manually on a 50-prompt set across four engines and you will see more signal than any rank report your current retainer produces. Add AI Growth Agent's automation and the loop runs at machine cadence without founder time.