Framework · Proposed standard

The AI Visibility Measurement Standard

'AI visibility' currently means whatever each vendor's dashboard happens to measure. A proposed open specification for what counts as a citation, how to sample, and how to report uncertainty — so numbers from different sources become comparable.

This solves a real problem: cross-vendor comparison is currently impossible because no two AI-visibility products define 'visibility' the same way. This is not a new metric — it's a specification other metrics can be checked against.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v1 proposal ·13 min read ·Last verified September 2026
The short version

Report four metrics separately — citation rate, share of voice, stability, cross-platform concordance — each with its sample, window, engine and citation definition attached. Never collapse them into a score out of 100. Follow that and any two measurements become comparable; skip it and "visibility" means whatever the reporter wants it to mean.

What this page establishes
  • Each metric below is written to be quoted verbatim: formula, unit, population and collection method in one block, so a definition can be lifted without the surrounding argument.
  • A properly reported figure names four things a score out of 100 hides — how many queries, on which engine, over what window, and what counted as a citation.
  • AI answers are non-deterministic, so a single-run measurement conflates signal with noise. This standard asks for three to five runs per query and the spread reported alongside the mean.
  • The standard deliberately excludes a composite visibility score. The weighting between inputs is the substantive claim, and it is the part vendors do not disclose.
  • This is a proposal, not an adopted standard. No vendor is bound to it, and most current tools do not follow it.

The four metrics, defined

Each block below is self-contained. Take one, quote it, and it carries its own formula, unit, population and collection method with no reference back to this page.

Citation Rate (CR)

Definition. The share of queries in a fixed set for which a given domain is cited at least once by a named engine.

Formula. CR = (queries in which the domain was cited) ÷ (total queries in the set).

Unit. A percentage, 0–100%.

Population. One named, frozen query set, measured on one named engine, within a stated collection window. Never a blend of engines.

Collection method. Every query run at least three times within the window. State in advance whether a query counts as cited when the domain appears in any run or in a majority of runs; the two produce different numbers on the same data.

What it solves. It is the base unit, and no two tools currently define it the same way.

AI Share of Voice (SOV)

Definition. A domain's citations as a proportion of all citations awarded across a fixed query set by a named engine.

Formula. SOV = (citations of the domain) ÷ (all citations of all domains in the same set and window).

Unit. A percentage, 0–100%.

Population. One named query set on one named engine. The denominator is engine-specific, so two SOV figures from two engines are not on the same scale.

Collection method. Count every cited domain in every run, not only your own. Report the composition of the denominator alongside the share, because a rise can mean a rival fell.

What it solves. It makes competitive comparison possible on a fixed denominator, which citation rate alone cannot.

Citation Stability

Definition. How consistently the same query returns the same cited sources across repeated identical runs.

Formula. For each query, the share of runs in which the domain appeared; stability for the set is the mean of those per-query shares across queries where the domain appeared at least once.

Unit. A proportion between 0 and 1.

Population. The same frozen query set, same engine, same window as the citation rate it accompanies.

Collection method. Three to five identical runs per query, issued from a controlled session — logged out, location fixed, no query history — because a measurement taken from a working browser is a measurement of that browser.

What it solves. It quantifies how much of a reported number is noise. A period-over-period movement smaller than your own run-to-run spread is not a result.

Cross-Platform Concordance

Definition. How much two engines agree on which domains to cite for the same query set.

Formula. A Jaccard index — the intersection of the two engines' cited-domain sets divided by their union.

Unit. A coefficient between 0 and 1; 1.0 is perfect agreement, 0 is no overlap at all.

Population. One query set, run on exactly two named engines, in the same window.

Collection method. Both engines sampled over the same dates with the same citation definition. A definition that differs between the two engines makes the coefficient meaningless.

What it solves. It prevents a result on one engine being reported as a result about AI search.

Concordance is in the set because published overlap figures are low and inconsistent.

Ahrefs reports limited overlap between engines; Seer found 87% of SearchGPT citations matching Bing's top results; and Search Engine Journal reports AI Overview citations drifting away from top-ranking pages. Three studies, three surfaces, three definitions. The gap between Google's own two surfaces is measured in the source delta study.

How to report a figure

Under this standard a citation-rate claim reads: "Domain X had a 34% citation rate across 200 queries on ChatGPT Search, collected across 5 runs each between [dates], with a stability score of 0.71." Not: "Domain X scored 82/100 on AI visibility." The first can be checked, disputed or replicated. The second cannot.

SampleHow many queries, and which set. Published if possible.
WindowThe dates collection started and ended. A number without one has no shelf life.
EngineNamed, singular. A blended cross-engine figure is a summary, not a measurement.
DefinitionWhat counted as a citation. Link, named mention, or paraphrase.
A score out of 100Unfalsifiable by construction. The weights are the claim, and they are the part not disclosed.
Before (typical vendor phrasing)After (standard-compliant)
"Your AI visibility score improved from 61 to 74 this quarter.""Citation rate across our 150-query brand set rose from 22% to 31% on ChatGPT Search (5 runs each, stability 0.68 → 0.74); Perplexity citation rate was flat at 18%. AI Overviews were not included in this quarter's measurement."

The compliant version is longer and less quotable as a headline. That trade-off is deliberate: it tells the reader exactly what improved, on which engine, with what confidence, and what was not measured at all. Applied to this site's own work, it is what the AI Citation Index reports and what the dataset strategy stores.

The five principles

Each principle gates the next
  1. 1 Define citation Link, named mention, or paraphrase. Written down before any counting starts.
  2. 2 Run repeatedly Three to five runs per query. Report the spread, not only the mean.
  3. 3 State sample and population How many queries, which engines, what window. Without this, nothing is falsifiable.
  4. 4 Label the claim Correlational or causal. Almost everything in this field is the former.
  5. 5 Publish components Citation rate, share of voice and stability separately. Never a single blended score.

Why the term is broken today

Fact

AI answers are non-deterministic. The same query can return a different set of cited sources on a second attempt, which is why concordance has to be measured across repeated runs rather than inferred from one. Almost every published AI-visibility figure comes from a single run.

Fact

"Citation" is not consistently defined. Whether an unlinked brand mention counts, or a paraphrase without attribution, is answered differently by different tools and usually not disclosed. The unlinked-mention question is not academic: including or excluding those mentions changes a citation rate by itself.

Fact

Sample and population are usually undisclosed. A visibility score rarely states how many queries it rests on, across which engines, over what window — making it unfalsifiable by construction. The same weakness runs through the figure set traced in the provenance audit.

The underlying reason is structural: every major player measuring this has a commercial reason to define it favourably. Even the careful vendor studies — Semrush on AI Overviews and Ahrefs on citations and the top 10 — sit on corpora nobody outside the company can inspect. Terms used here without definition are in the glossary.

Building a compliant query set

Every principle depends on the query set, and the query set is where most measurement quietly fails. A biased set produces disclosed, checkable, useless numbers.

Start from demand, not from your content. Queries chosen because you already rank for them guarantee a flattering result, and a set built from pages you already have will keep rediscovering the domains that already dominate.

Include queries you expect to lose. A set with no failures is not measuring anything. Losses are where change becomes visible first.

Cover the intent range deliberately. Definitional, comparison, how-to, purchase. Intent matters because engines fan a single question out into several sub-searches, as the granted patent describes and the fan-out corpus study measures.

Fix the exact wording and freeze it. Paraphrases are different queries. A set that drifts cannot support a trend claim.

Hold back a control group. Queries relevant to your category that you are not working on. When everything moves together, they tell you the platform moved, not you.

Publish the set if you can. A published query set is the single strongest disclosure available. It lets someone else run your measurement and disagree with your conclusion using your own instrument.

Computing stability, with arithmetic

The numbers here are invented for illustration. They are not measurements.

Run one query five times and your domain is cited in three of them: for that query you were cited 60% of the time. Run a second query five times and it is cited in all five. A single-run tool records both as "cited". They are not the same result, and the difference is exactly what stability captures.

IllustrativeA set where most cited queries are cited in five of five runs is a stable set. A reported change is likely real.
IllustrativeA set where most cited queries appear in two or three of five is unstable. A period-over-period change of a few points is probably noise.
The rule that followsNever report a movement smaller than your own run-to-run spread as a result.

That is the practical payoff of principle two. It converts "did it change" from a comparison of two numbers into a comparison against noise. Without it, every measurement programme reports movement every period, because something always moves.

Seven pitfalls that break comparability

One: changing the query set between periods. The most common and most damaging. A comparison across two different sets is not a trend.

Two: silently changing the citation definition. Counting unlinked brand mentions this quarter and not last quarter produces a rise on its own, and correlational ranking-factor work suggests that is a large swing rather than a rounding difference.

Three: mixing engines into one figure. A blended rate hides which engine moved, and engines behave very differently — see this side-by-side of how each platform cites.

Four: counting domain instead of URL, or the reverse. Both are defensible. Switching between them mid-programme is not.

Five: uncontrolled personalisation. Signed-in sessions, saved location, prior queries. Google's documentation on AI features is the only first-party account of what varies, and it does not describe personalisation in enough detail to control for it.

Six: sampling at a fixed time of day only. If system behaviour varies with load, a fixed sampling time bakes that variation into your trend.

Seven: reporting the mean without the spread. A mean alone cannot be interpreted. This is principle two restated, because it is the one most often skipped under deadline.

What does not transfer between engines

Does not transferWhy not
Citation rateEngines differ in how many sources they cite at all. A higher rate can be an interface property, not a visibility gain.
StabilitySome engines are far more repeatable than others. Comparing raw stability across engines compares their architectures.
Share of voiceThe denominator is engine-specific. Two share figures from two engines are not on the same scale.
The effect of a changeA page edit visible in one engine within days may take much longer to register in another, or never — the question the crawl-to-citation latency study exists to measure.
What counts as a citationInline links, source lists and unlinked mentions appear in different proportions per engine.

The standard's answer is not to normalise these away. It is to report per engine, always, and to treat any cross-engine aggregate as a summary rather than a measurement.

Why no composite score

The exclusion of a composite score is the most contested part of this proposal. It rests on four points.

The weights are the claim, and they are hidden. A score combining citation rate, share of voice and sentiment embeds a theory about how much each matters, and that theory is exactly the part not disclosed. Compensation hides direction: a score can hold steady while citation rate falls and something else rises.

It cannot be replicated. Every component in this standard can be recomputed by an outsider from disclosed inputs; a proprietary blend cannot. And it becomes a target — a score clients are managed against will be optimised, including by the vendor who defines it. This is the same objection the tactic evidence scoreboard makes to blended tactic rankings.

Hypothesis

We expect composite scores to persist regardless of this argument, because they are commercially useful. The realistic goal is disclosure of components alongside the score, not the score's disappearance.

Confounds in period-over-period reporting

Even a fully compliant measurement can support a wrong conclusion.

The platform changed, not you. Engines ship changes without announcement. Control queries are the only cheap defence, and almost nobody keeps them.

Competitors changed. Share of voice is relative; a rise can mean a rival got worse rather than you got better.

Seasonality in the query set. Some questions are asked differently at different times of year, so a quarter-over-quarter comparison can be a comparison of seasons.

Your own unrelated changes. A migration, a redesign or a content pruning will move these numbers. Keep a change log alongside the measurement log or attribution becomes guesswork.

Regression to the mean. A period measured after an unusually bad one tends to look better regardless of action taken. This is the confound most likely to make an intervention look effective when it was not.

The clearest published example of what a compliant negative result looks like is the llms.txt evidence: an adoption study, a request-tracking study, and a 300,000-domain analysis finding no clear effect. This site's own version is a first-party log study.

A minimum viable version

Full compliance costs real money. A stripped version is far better than nothing, and it is honest about being stripped.

Thirty queries, frozen, five of them controls. Small, but enough to see large movements and to catch platform-wide shifts.

Three runs per query, not five. Three is enough to distinguish "always" from "sometimes", which is most of the value.

One engine, named. Measuring one engine properly beats measuring four badly. Check first that the engine can reach you at all — access gates everything, which is what the bot registry and the technical audit cover.

One definition of citation, written down. Whatever you choose is fine as long as it is fixed and stated.

One sentence of disclosure attached to every figure. Sample, window, engine, definition. That single sentence is the standard's core.

This version is cheap enough for a small team to run monthly. It is not as good as the full specification, and saying so is part of complying with it.

What would move this to v2

What would move this specification to v2
  1. Evidence on stability thresholdsAwaiting data

    Principle 2

    We decline to publish a threshold because we have no basis for one. Data on typical run-to-run spread would let a real threshold replace judgement.

  2. Operator sampling interfacesNot offered

    All engines

    If engines exposed a controlled sampling interface, most of the sampling difficulty here would disappear.

  3. Evidence a component is redundantOpen

    Metric set

    If concordance carried no information beyond the per-engine rates, we would drop it. A standard that only ever grows is one nobody is testing.

Limitations

  • This is a proposal, not an adopted standard. No vendor is bound to it, and most current tools do not follow it.
  • Perfect compliance is expensive. Multi-run sampling across several engines costs more than a single-run check; the standard states the ideal, not a requirement for every casual use case.
  • It does not resolve the definition of "citation" itself — it requires you to state your definition, not adopt a specific one, because reasonable definitions vary by use case.
  • Disclosure alone does not prevent a poorly designed sample from producing a misleading, if technically checkable, result.
  • Citation rate is an intermediate metric. Whether it produces revenue is a separate question this standard does not address — see conversion benchmarks.
  • Not every surface can be sampled repeatably. Where principle two is unachievable, say so rather than reporting a single-run number as compliant.

Last verified: September 2026

September 2026
  • Rewrote the four metrics as self-contained definition blocks carrying formula, unit, population and collection method, so each can be quoted without this page.
  • Cut the media-measurement analogy, the objections section and the stakeholder-reporting section; none carried evidence.
  • Moved the metric definitions above the diagnosis of why the term is broken.
Earlier
  • v1 proposal published: five principles, four metrics, no composite score.

Next: run the minimum viable version on your own site. Start by checking which questions you currently appear on with the AI Overview checker, then freeze that list as your query set. To see the full specification applied end to end, read the AI Citation Index.

How to cite this
Namdev, R. (2026). The AI Visibility Measurement Standard (v3). Retrieved from https://ritiknamdev.com/blog/ai-visibility-measurement-standard

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

This is the measurement pillar under AI SEO. It is applied directly in the AI Citation Index's methodology, referenced by the GEO tactic evidence scoreboard's grading system, and used to report the entries in the null-results registry.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

GEO: Generative Engine Optimization — Aggarwal et al., KDD 2024arxiv.org/abs/2311.09735 arXiv — GEO paper, full PDF (benchmark design and metric definition)arxiv.org/pdf/2311.09735 Ahrefs — AI search and Google ranking overlapahrefs.com/blog/ai-search-overlap Ahrefs — AI Overview citations and the top 10ahrefs.com/blog/ai-overview-citations-top-10 Ahrefs — AI Overviews reduce clicks (updated study)ahrefs.com/blog/ai-overviews-reduce-clicks-update Ahrefs — AI SEO statisticsahrefs.com/blog/ai-seo-statistics Semrush — AI Overviews studywww.semrush.com/blog/semrush-ai-overviews-study Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Search Engine Journal — AI Overview citations from top-ranking pages drop sharplywww.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637 Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Discovered Labs — How ChatGPT, Claude and Perplexity choose sourcesdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Discovered Labs — How each platform cites sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Leapd — How ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Ziptie — How does ChatGPT choose its sources?ziptie.dev/blog/how-does-chatgpt-choose-its-sources Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors Similarweb — Generative AI usage statisticsaisearch.similarweb.com/blog/gen-ai-stats Statista — Global monthly ChatGPT userswww.statista.com/statistics/1659718/global-monthly-chatgpt-users StatCounter — Search engine market sharegs.statcounter.com/search-engine-market-share Search Engine Journal — llms.txt shows no clear effect on AI citations across 300k domainswww.searchenginejournal.com/llms-txt-shows-no-clear-effect-on-ai-citations-based-on-300k-domains/561542 Ahrefs — llms.txt adoption studyahrefs.com/blog/llmstxt-study Originality.ai — llms.txt tracking studyoriginality.ai/blog/llms-txt-tracking-study Google Patents — US11663201B2, query fan-out and sub-query categoriespatents.google.com/patent/US11663201B2 Google Search Central — AI features and your websitedevelopers.google.com/search/docs/appearance/ai-features ritiknamdev.com — Where AI SEO statistics come fromritiknamdev.com/blog/where-ai-seo-statistics-come-from
FAQ

Frequently asked questions

Is this an official industry standard?
No. There is no standards body for AI-search measurement, which is exactly the gap this page addresses. It is a proposed specification, published openly so anyone can adopt, critique, or improve it.
Why not just use whatever metric my visibility tool reports?
You can, but know what you are getting. Most vendor dashboards report a single-run snapshot as if it were a stable measurement, and many blend multiple signals into one opaque score. This standard asks for the raw ingredients - citation rate, sample size, variance - so you can judge the number yourself.
Why exclude a composite visibility score?
Because it is unfalsifiable. A composite score lets a vendor choose which inputs to weight and how, with no way for an outside observer to check the math. Every underlying component this standard asks for can be independently verified. A single blended number cannot.
How many queries does a query set need?
Enough that adding more stops changing the answer, which is a property of your queries rather than a fixed number. The practical test is to compute your citation rate on a random half of the set, then on the other half. If the two halves disagree substantially, the set is too small or too varied to support the precision you are reporting.
What is a good stability score?
We are deliberately not publishing a threshold, because we have no evidence for where one should sit. A number invented to sound authoritative would be exactly the failure this standard exists to prevent. What stability is for is comparison: between your own periods, between engines, and between your query set and someone else's. A change in citation rate smaller than your own run-to-run spread is not a finding.
Should the same query set be used forever?
Mostly yes, and changing it is the most common way a measurement programme quietly breaks. A query set that evolves alongside your content will always show improvement, because you are changing the test and the answer together. Freeze a core set for trend reporting, and add new queries as a separate labelled set.
Does this standard tell me what to do to improve visibility?
No, and that is intentional. It is a measurement specification, not a tactic guide. Its purpose is to make claims about tactics checkable, including our own. A standard that also recommended tactics would have an interest in those tactics measuring well.
Could a vendor adopt this standard and still mislead people?
Less easily than today, but not impossible. A vendor could disclose a small, cherry-picked sample honestly and still produce a misleading impression. The standard makes a claim checkable. It does not guarantee every checkable claim is well-designed. Checkability is a necessary condition for trust, not a sufficient one.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.