AI visibility

How to measure AI share of voice for an ecommerce store

Shoppers now ask assistants what to buy. This guide explains what AI share of voice actually measures, how to probe for it without fooling yourself, which metrics are worth a dashboard, and what to do when a prompt comes back without your name in it.

Last updated August 2026 · 7 min read

What AI share of voice actually is

Ask an assistant for the best running shoes for flat feet and it will name a few brands. Maybe four. Maybe seven. Your store is either in that list or it isn't.

AI share of voice is the answer to a narrow question: across all the brands an assistant names in response to your tracked shopper prompts, what proportion of those mentions are yours? If you run ten probes and the assistants name thirty brands in total across those answers, and six of the thirty are you, your share of voice is 20%. That is the whole calculation. The difficulty is never the arithmetic. It is getting probes that mean something.

Note what the metric does not measure. It does not measure how many people asked. It does not measure clicks. It measures presence in the recommendation set — the shortlist an assistant hands a shopper who is deciding what to buy.

The surface is big enough to be worth instrumenting. Perplexity alone handled 780 million queries in May 2025, a figure given by CEO Aravind Srinivas alongside month-over-month growth of more than 20% at that time (TechCrunch, 5 June 2025). That is one assistant in one month, and no newer official number has been disclosed since — which is itself a useful reminder that you are measuring a channel nobody reports on consistently. Your own probe data will often be better evidence than anything published about it.

Why a rank-tracking mindset doesn't transfer

If you came from SEO, the instinct is to treat this like a rank tracker: check the position, watch it move, celebrate page one. That instinct will mislead you in three specific ways.

First, there is no single result page. Two shoppers asking near-identical questions can get different brands, in a different order, with different reasoning. There is no shared artifact to rank within.

Second, the output is unstable by design. A rank of 4 today usually means a rank near 4 tomorrow. An AI mention today means only that a mention was likely enough to happen once.

Third, the unit is the brand, not the URL. Assistants synthesize. They name a store, sometimes a specific product, and they may or may not cite a source you can attribute. Optimizing a single landing page for a single phrase is not the lever it used to be.

THE SHIFT

Stop thinking "what position do I hold" and start thinking "how often do I show up." AI share of voice is a frequency estimate, not a coordinate. Every conclusion you draw should be phrased in those terms.

The metrics worth tracking

Four numbers carry most of the signal. Track all four, because each one fails to tell you something the others cover.

MetricWhat it tells youHow it's calculated / what good looks like
Mention rateWhether the assistant knows you exist for this kind of request at allProbes where your brand appears, divided by total probes. Mentioned in 3 of 10 probes is a 30% mention rate. Zero on a core category prompt is the loudest signal on your dashboard.
Average positionHow prominently you're placed once you are mentionedYour ordinal spot in the list of brands named, averaged across the probes where you appeared. Earlier tends to carry more weight, since shoppers rarely work down to the seventh suggestion.
Citation rateWhether the assistant treats your site as a source, not just a name it recallsProbes that link or cite your domain, divided by probes where you were mentioned. A high mention rate with a low citation rate means you're remembered but not read.
Share of voiceYour presence relative to a named competitor setYour mentions divided by all brand mentions across the same probe set. Best read as a trend against the same competitors and the same prompts over time.

Read them together. A rising mention rate with a flat position means you are entering more conversations but not winning them. A flat mention rate with a falling share of voice means competitors are being added to lists you were already on.

Citation rate earns its own column because the behaviour underneath it is still moving. Similarweb's tracking put ChatGPT's citation presence in US prompts at roughly 1.6% in June 2025, rising to 6.8% by May 2026 (Similarweb, 29 July 2026). When how often assistants cite anything at all is itself changing year on year, your own citation rate has to be read as a trend against a shifting baseline, not as a score you can compare to somebody else's benchmark.

Non-determinism, and what honest methodology looks like

This is the part most people skip, and it is the part that decides whether your data is worth anything.

The same prompt, sent to the same assistant, twice in a row, can produce different brands. Not occasionally — routinely. Sampling variation, model updates, and retrieval freshness all move the output. Nobody controls this, and no tool eliminates it.

Which means a single check is close to meaningless. If you probe once and see your name, you have not proven you are visible. If you probe once and don't, you have not proven you are invisible. You have one sample from a distribution.

WHY YOUR OWN METHOD BEATS A HEADLINE NUMBER

Take a question vendors have actually tried to answer: how much do AI citations overlap with classic search rankings? BrightEdge, reporting on a 16-month study across nine industries, stated that "Only 16.7% of citations come from top 10 results" (BrightEdge, 18 September 2025). Other vendors measuring what they describe as the same overlap have published figures as high as 99.5%. That spread — roughly 16.7% to 99.5% — is not a rounding disagreement. It is different tools, different prompt sets, different time periods, and different ideas of what counts as a citation. Treat none of the published figures as settled, including the one quoted here. It is the plainest argument available for defining your own measurement and holding it constant, rather than importing a number you cannot reproduce.

Sound methodology follows from that:

  • Repeat probes. One run is an anecdote. Repetition on a schedule turns it into an estimate.
  • Read the trend, not the snapshot. Compare this month to last month, not Tuesday to Wednesday.
  • Hold the prompt set constant. Rewording a tracked prompt breaks its history. If you must change it, treat it as a new line.
  • Break results out per engine. ChatGPT, Perplexity, Gemini, Claude and Google AI draw on different sources with different freshness. A blended average hides where you're actually losing.
  • State the number as an estimate. "Roughly a third of probes" is honest. "We rank 30%" is not.

The per-engine point is not housekeeping. Similarweb's tracking shows ChatGPT's share of generative-AI website visits falling from about 76% in June 2025 to roughly 53% by May 2026, while Gemini rose from under 9% to around 27–28% and Claude from barely 2% to close to 9% (Similarweb, 29 July 2026). A blended average assembled when one assistant dominated will quietly misrepresent your position a year later, and a prompt set tuned to a single engine's habits is a bet on a distribution that has already shifted once. Track several, and expect the weighting between them to keep moving.

Cuebase's Share of Voice module is built around this: it sends real shopper prompts to the assistants on a fixed cadence and stores the results so you get trend lines per query and per engine rather than a one-off reading. It's a paid feature, starting on the Starter plan — it isn't included on Free.

Choosing prompts worth tracking

Prompt selection matters far more than prompt count. Twenty prompts that mirror how your buyers actually talk will beat two hundred generic ones.

Real shopper prompts are specific. They stack an attribute, a use case, and a constraint. They are the sentence someone types when they're close to spending money.

Weak, too broad: best hiking boots

Strong, closer to how people ask:

  • best waterproof hiking boots for wide feet under $200
  • what's a good espresso machine for a small apartment kitchen
  • compare merino base layers for cold weather cycling
  • gift ideas for someone getting into film photography under $150
  • durable dog harness for a puller that won't chafe

Spread your set across intents rather than piling into one:

  • Discovery — "what should I look for in X"
  • Comparison — "X vs Y for Z"
  • Budget — "best X under $N"
  • Use case — "X for [specific situation]"
  • Gift — "gift for someone who [does thing]"

Then leave them alone. The value compounds with time held constant, and every edit resets the clock on that line.

Setting up measurement, step by step

  1. Pick your prompt set

    Start from your top categories and your actual customer language — support tickets, on-site search, product reviews. Write prompts a shopper would plausibly type. Cover discovery, comparison, budget, use case and gift intents.

  2. Name your competitor set

    List the brands you genuinely lose deals to, not the aspirational ones. Share of voice is only interpretable against a fixed comparison group, so pick it deliberately and keep it stable.

  3. Set a cadence and hold it

    Weekly is enough to see movement without drowning in noise. Faster cadences give you more samples per period, which tightens the estimate. Whatever you choose, keep the spacing even.

  4. Break results out by engine

    Store every result tagged by engine and by query. You need per-engine views because the engines disagree, and a single blended number will hide the disagreement.

  5. Read the trend, not the last run

    Give it several cycles before drawing conclusions. Look for sustained direction across multiple probes on the same prompt. Ignore single-run swings — they're expected.

  6. Act on the persistent gaps

    Find the prompts where you're consistently absent, identify who's being named instead, and work backwards to why. That's your queue.

AI visibility is not AI traffic

These are different questions and they need different instruments.

An assistant can recommend your store and send you nothing. The shopper copies the brand name and searches it directly. Or reads the recommendation, closes the tab, and buys three days later. Or completes the purchase inside the chat surface entirely. In every one of those cases you were visible, and your referral data will not show it.

The reverse also happens: a click arrives from an AI surface for a prompt you never tracked.

The gap is measurable, not just plausible. Pew Research Center, using the browsing data of 900 US adults who agreed to share it in March 2025, found users clicked a traditional result link on 8% of visits where an AI summary appeared, against 15% of visits where none did — and just 1% of visits to a page carrying an AI summary produced a click on a cited source (Pew Research Center, 22 July 2025). That study covers Google's AI summaries rather than a chat assistant, and it is one panel at one point in time, so do not read it as a conversion rate for your own store. Read it as evidence that the mechanism is real: being named in the answer and receiving no click are entirely compatible outcomes.

So read both. Mention data tells you whether assistants are recommending you. Pixel and analytics data tell you what happens to the shoppers who do click through. Cuebase pairs the two — Share of Voice for the mention side, Agent Channel Analytics for AI-referred traffic and conversion broken down by source — but the principle holds regardless of tooling. Neither number alone describes your position in AI-mediated shopping.

Turning a gap into a fix

A persistent zero on a prompt is a lead, not a verdict. It usually traces back to one of two things.

The first is your own product data. If your listings for that category are thin, ambiguous, or missing the attributes the prompt asks about — the width, the capacity, the material, the use case — there is nothing for a model to match against. "Waterproof hiking boots for wide feet" needs your catalog to actually say waterproof, and actually say wide.

The second is a competitor who has covered that use case more clearly. Their content answers the question in the shape the question was asked. That's a content gap you can close.

Either way the work is the same: make the product data unambiguous for the specific intent you're losing. This is what Cuebase's Catalog Readiness Hub is for — it scores each product against an agent-readiness rubric, flags the gaps, generates fix suggestions, and applies them to Shopify once you've reviewed them.

Then wait. Measure again in a few cycles. And expect movement to be gradual and noisy rather than clean — that's the nature of the surface, and no amount of optimization makes an assistant's output guaranteed.

FAQ

How often should I check my AI share of voice?

On a schedule, not on impulse. A fixed cadence such as weekly, every three days, or every two days gives you evenly spaced points on a trend line, which is what makes movement readable. Checking manually whenever you feel curious produces uneven spacing and tempts you to react to noise. Pick a cadence you can hold, and read the data at a slower rhythm than you collect it.

Why did my score change when I did not change anything?

Because AI answers are non-deterministic. The same prompt sent twice can return a different list of brands. Model updates, retrieval freshness, and ordinary sampling variation all move the number without anything changing on your store. This is expected. Judge your visibility by the direction of the trend across many probes, not by any single run.

Is AI share of voice the same as a keyword ranking?

No. A ranking is a stable, ordered position in a list of links that most people looking at the same query would see. AI share of voice is a probability estimate: how likely an assistant is to name your brand among the handful it mentions in a synthesized answer. There is no position 1 to hold, and there is no single answer everyone sees.

How many prompts should I track?

Fewer than you think, chosen more carefully than you think. A tight set of prompts that mirror how your actual buyers ask, held constant over months, will teach you more than a long list you keep editing. Cover your main categories and a spread of intents, then stop adding. Every prompt you change resets the history for that line.

Is Share of Voice included in the Cuebase free plan?

No. The Free plan covers catalog readiness scoring, AI channel analytics, and AI fix suggestions. Share of Voice starts on Starter at 29 dollars per month with 10 tracked queries across 3 engines on a weekly cadence. Growth is 79 dollars per month with 20 queries across 5 engines every 3 days plus competitor tracking, and Pro is 199 dollars per month with 40 queries across 5 engines every 2 days.


Sources

Every figure quoted above is linked to its original publisher with the date attached, so you can check it and judge how current it is. Where the published measurements disagree — and on citation overlap they disagree enormously — we have shown the spread rather than picking the number that reads best.

  1. BrightEdge. Rank overlap after 16 months of AIO. Weekly AI Search Insights, 18 September 2025. Quoted here as one vendor measurement among several that disagree, not as a settled figure.
  2. Pew Research Center. Google users are less likely to click on links when an AI summary appears in the results. 22 July 2025. Based on the browsing data of 900 US adults, March 2025.
  3. Similarweb. Generative AI usage statistics. 29 July 2026.
  4. TechCrunch. Perplexity received 780 million queries last month, CEO says. 5 June 2025. The figure covers May 2025 only; no later official number has been published.

Measure where you stand in AI answers

Cuebase is the AI SEO and GEO layer for Shopify: catalog readiness scoring, AI channel analytics, and Share of Voice tracking across ChatGPT, Perplexity, Gemini, Claude and Google AI. Free plan available; Share of Voice starts on Starter.

Add to Shopify →