What is a good AI visibility score? What the number measures and where it stops being useful

What is a good AI visibility score? What the number measures and where it stops being useful

In short

Every tool in this market returns a number called an AI visibility score, and no two of them compute the same thing. This page separates the two families these scores fall into, gives you a usable reading for each, and says plainly what the number does not tell you. scaile publishes the arithmetic behind its own score for the same reason: a diagnostic you cannot check is a diagnostic you cannot act on.

Published
Reading time
10 min
Last updated

Simon is co-founder of scaile and works with enterprise teams on visibility in AI search, from the first audit to defensible Share of Voice numbers.

Key takeaways

  • No standard exists. Semrush returns a 0 to 100 benchmark, Peec and Profound both return percentages but divide by different denominators, and Ahrefs Brand Radar returns no single score at all. Four vendors, four numbers, none of them comparable with another.

  • Scores fall into two families. Readiness scores are computed from your own site and are deterministic in what they ask: the same checks over the same evidence give the same score. Presence scores are computed from sampled model answers and move on their own.

  • The movement is not a rounding error. Answers are assembled live from retrieval, and AI search results are, in Semrush's own words, fast-changing and highly personalised — so a presence number sampled once is a single draw, not a level.

  • scaile's free Health Check uses a catalogue of 103 checks, a score from 0 to 100, a letter grade, and a stated count of how many checks actually ran. The AI Visibility Check starts with one buying question; the requested report extends the sample to ten.

  • A readiness score is not a ranking and does not predict traffic. It tests access and machine readability; it does not prove that a search engine has indexed a page or will cite it.

There is no industry-standard AI visibility score. On scaile’s Health Check, 90 and above is labelled Good, 70 to 89 Needs Work, and below 70 Critical. These are our diagnostic bands, not industry benchmarks or proof that every important check passed. Read the findings and the number of checks that actually ran. For a presence measurement, compare your brand with the same competitors on the same questions, engines and dates; there is no universal good percentage.

Anyone who quotes you a benchmark without naming the instrument is quoting noise.

Both numbers describe a gap; neither closes one. Technical fixes can improve readiness. Useful, sourced content and credible third-party coverage can improve the material an assistant can retrieve, without guaranteeing a particular presence score. scaile runs the content-production part as managed infrastructure: a named AI Search Strategist, research from your knowledge base, fact-checking and your team’s approval before publication.

Why is there no standard AI visibility score?

Because every vendor invented its own, and the definitions genuinely differ.

Semrush’s AI Visibility Score is documented as “a benchmark score (0–100) showing how often your brand appears in AI-generated answers compared to competitors.” It is not built by asking models your questions: Semrush draws on a database of over 317 million prompts and responses across ChatGPT, Gemini, AI Overviews and AI Mode, captured from clickstream data rather than through model APIs, covering 117 regional databases. The prompt database refreshes daily on a rolling basis; the brand reports built on it refresh weekly, so a dashboard you open on Tuesday is not showing you yesterday. The exact weighting is proprietary.

Peec defines its Visibility Score with an explicit formula: responses mentioning your brand, divided by total responses, times 100. Profound also reports a Visibility Score as a percentage, but divides brand appearances by the total responses containing at least one brand — a different denominator, from prompts Profound sends to answer engines daily. Run the same brand through both on the same day with the same questions and the two numbers will disagree, and both will be correct.

Ahrefs took a fourth road. Brand Radar reports brand mentions across the major AI platforms — the list has grown twice in 2026, so check the product page for what is covered today — plus an AI Share of Voice against competitors, with no 0 to 100 score at all. Its own guidance says AI platforms are non-deterministic and that you should track the range of your share of voice rather than a point value.

Two things follow. First, “our AI visibility score is 62” is not a fact about your company until you say whose 62 it is. Second, no industry benchmark exists to compare it against, because a benchmark requires an instrument everyone shares. Semrush says so itself in its own documentation: AI search responses are fast-changing and highly personalised, “which means no platform can provide exact numbers.”

Even the research had to invent its own yardstick. The GEO paper that established this field measured visibility with three metrics it defined from scratch — a normalised word count, a position-adjusted word count that decays exponentially with where the citation appears, and a subjective impression score scored by a language model — precisely because a ranking position, the metric classic search runs on, does not describe a generative answer.

What are the two families of AI visibility score?

Every score on the market is one of two things, and confusing them is the most expensive mistake in this whole topic.

Readiness scores are computed from your own site. Can the AI crawlers reach you? Can a machine parse an answer out of the page? Does the page carry a date, an author and a source, so the answer can attribute it? These are questions about files that exist on your server right now. A readiness score is deterministic in what it asks, not in what it can reach: the same checks over the same evidence give the same score, and the ran-count the report prints is what tells you whether the evidence was the same. A third-party API that times out, a page that would not render, a robots.txt that came back a 503 — each one takes a check out of the denominator, and a score built on 62 checks is not the same measurement as one built on 95.

Presence scores are computed from sampled model answers. Ask an assistant a question, read the answer, record whether you were named. Repeat. These measure something readiness cannot: whether you actually come up when a buyer asks. They also move on their own, and more than most dashboards admit.

How do you calculate AI visibility? A worked example

This is an illustrative calculation, not a client result or industry benchmark. Suppose you ask 20 non-branded buying questions across three engines on two dates and receive all 120 answers.

MeasurementCalculationResult
Usable answers20 questions × 3 engines × 2 runs120
Answers naming your brand24 ÷ 12020% mention rate
Answers citing your website12 ÷ 12010% own-site citation rate
Next comparable sample naming your brand36 ÷ 12030% mention rate
Change in mention rate30% − 20%10 percentage points, or 50% relative growth

Count an answer once for each metric, even if it repeats the brand or links to the website several times. A mention without a link counts toward mentions, not website citations. Keep failed requests separate: 120 attempts are not 120 usable answers if an engine times out. If an engine does not expose sources, label its citation data unavailable instead of silently treating it as zero citations.

Keep a log with the question, engine, date, language, market, brand mentioned, own website cited and source URLs. Repeat the same question set before calling a change progress. The 120 answers are not necessarily independent observations, and this calculation alone does not establish statistical significance or predict traffic.

This simple mention rate is not scaile’s composite AI Visibility Check score. That score combines presence, rank, share of voice and citations; compare like with like. Start with the free AI Visibility Check, then use a fixed repeated sample for ongoing measurement.

Why does a presence score move on its own?

Because the answer is assembled live, and almost nothing in that assembly is held fixed between two runs.

The dominant cause is retrieval, not the model. Semrush states it plainly in its own documentation: AI search responses are fast-changing and highly personalised, which is why “no platform can provide exact numbers.” Ahrefs gives the same warning from the other side, telling Brand Radar users that AI platforms are non-deterministic and that the honest reading is a range rather than a point value. Between your Monday sample and your Thursday sample, the search index behind the assistant has churned, the pages it retrieved have changed, and the personalisation layer has treated the two requests differently.

Underneath that sits a smaller, more surprising floor. Thinking Machines sampled the same prompt 1,000 times at temperature zero — the setting that is supposed to make output deterministic — on an open-weights model and got 80 distinct completions, the most common occurring 78 times, with the first divergence at the 103rd token. The cause was not randomness in the sampler but batch-size variation on the serving hardware, which changes numerical results depending on how many other people happen to be querying at the same moment. The important half of that result is the fix: with batch-invariant kernels, all 1,000 completions came back identical. So this is not evidence of irreducible randomness in language models. It is evidence that even the floor case is not reproducible on ordinary serving infrastructure, before retrieval and personalisation add their much larger movement on top.

The GEO researchers handled the same problem the honest way: they sampled five responses per query at temperature 0.7 “to reduce statistical deviations,” on GEO-bench, a benchmark of 10,000 queries. That is what a presence measurement costs if you want it to hold still.

So when a dashboard shows your presence score moving from 34 percent to 31 percent week on week, the first question is not what caused the drop. It is whether three points is even outside the instrument’s own noise floor. On ten questions asked once each, one answer flipping moves the number ten points. Nothing about your website changed.

This is why scaile’s AI Visibility Check hands you the underlying answers rather than only a headline number: what the assistants actually said is checkable, and a percentage derived from a single pass is not. Continuous tracking with enough repetition to separate signal from noise is a different job, and it is what the reporting side of the platform does.

What counts as a good readiness score?

This is the family where a number in the abstract does mean something, because the same checks over the same evidence always produce the same score.

On the 0 to 100 scale the free AI Search Health Check uses:

90 and above is Good: A from 90 to 94, A+ from 95. It is a strong result on the evidence checked, not a guarantee that no high-severity finding or untested technical problem remains. Review individual findings and coverage even when the headline is green.

70 to 89 is Needs Work: C from 70 to 79, B from 80 to 89. Prioritize the named findings by severity and affected pages. The total alone does not tell you whether a fix is small or whether a particular crawler is blocked.

Below 70 is where triage beats a to-do list. Read the result top-down by severity rather than working down the list, because a single problem in the access layer can account for a whole cluster of failures underneath it.

The distinction that matters at that layer is which bot you blocked. Blocking the retrieval agents — OAI-SearchBot, PerplexityBot — is the costly block: those are the agents that fetch a page at answer time, and in scaile’s own engine a robots.txt that disallows the retrieval crawlers together trips a score gate that caps the result outright, no matter what else passes. Blocking the training crawlers — GPTBot, ClaudeBot, CCBot — is a rights decision about your content being used to train future models. It is a legitimate one to make, and it barely moves whether ChatGPT or Perplexity cite you today, because citation runs through the retrieval agents. Fix the gate and the cap lifts; fix a training-crawler line and you have changed your licensing posture, not your visibility.

A score is only trustworthy if it shows its arithmetic. The Health Check reports how many checks passed out of how many actually ran, not out of 103, because the free pass reads your homepage, your robots.txt and your sitemap — enough to settle around 60 of the checks — and the rest abstain rather than guess. A check that cannot reach a verdict should lower your confidence in the score, not silently lower the score.

What does the score not tell you?

A readiness score is not a ranking, and it does not predict traffic. It is a precondition test.

A readiness score is not calibrated to forecast visits or revenue. Google includes AI Overviews and AI Mode in overall Search Console web-search reporting, rather than a separate AI-features report. A mention with no click creates no visit in website analytics, and referral attribution is incomplete. You can still measure observable visits and leads and run experiments; those limitations do not make research impossible, but they do make a score-to-revenue guarantee unjustified.

Readiness checks can identify access problems, but they do not establish Google’s indexing decision. A supporting page in AI Overviews or AI Mode must be indexed and snippet-eligible. Normal Search policies and helpful-content guidance still apply; special AI files or schema are not required, and eligibility does not guarantee selection. Confirm the stored indexing state separately in Search Console. When several access checks fail together, inspect shared causes such as CDN rules, DNS failures or timeouts before treating each failure as a separate content defect.

You can score 95 and still be absent from the sampled answers. Investigate indexing, relevance to the question, useful evidence, third-party coverage, competing sources and sample variation. Content gaps are one possible cause, not a diagnosis established by the score alone. Use customer-intent research to identify specific unanswered buying questions, then check the other causes of invisibility before choosing an intervention.

How should you use the two numbers together?

Run a readiness check to identify access and parsing problems, then measure presence separately. A blocked crawler can prevent direct retrieval of your pages, but an assistant may still mention your brand from third-party sources. A readiness score therefore cannot replace a presence measurement, and presence does not prove your own pages are accessible.

Then measure presence against a fixed question set and named competitors. Record how often, where and from which sources the brand appears. Repeat the sample before interpreting a change. A low result is a reason to investigate; a readiness score alone cannot distinguish all the possible causes.

Both scaile checks are free and neither needs an account. The Health Check shows a readiness result on the page. The AI Visibility Check first asks one buying question across the available assistants and shows the result. The full report expands to ten questions with competitor and source analysis; requesting it requires an email address. The Sitemap Visualiser provides a separate structural view.

FAQ

Is 70 a good AI visibility score?

On scaile’s Health Check, 70 is Needs Work, grade C. Read the underlying findings rather than treating it as clearance. A 70% mention rate means the brand appeared in seven out of ten usable answers; whether that is good depends on the question set, competitors and measurement method. The two figures are not comparable.

Can I compare my Semrush AI visibility score with another tool’s?

No. Semrush computes a 0 to 100 benchmark from a clickstream prompt database, Peec divides brand mentions by all responses, and Profound divides by responses containing at least one brand. Different inputs, different denominators, different numbers for the same brand on the same day.

Why does my AI visibility score change when nothing changed on my site?

Because presence scores sample live answers, and those answers are rebuilt from live retrieval on every run, over an index that churns, with personalisation on top. Semrush and Ahrefs both say so in their own documentation. A readiness score computed from your own files does not move unless your files, or what the checker could reach, do.

Does a high AI visibility score mean more traffic?

No. A readiness score describes checks on your site, and a presence score describes sampled answers. Neither guarantees clicks. Track observable referrals, qualified leads and repeated visibility measurements separately, and state where attribution is unavailable.

How is scaile’s AI visibility score calculated?

There are two different scores. The Health Check uses a catalogue of 103 checks, severity-weighted deductions, partial results and explicit score caps; unverified checks do not count as failures. The AI Visibility Check’s composite weights presence at 40%, rank at 20%, share of voice at 20% and citations at 20%. Its report explains the inputs and unavailable source data. Neither is an industry-standard ranking.

Should I track an AI visibility score every week?

Only if the measurement repeats enough times to separate a real move from noise. A weekly single pass over ten questions is a coin-flip dressed as a trend line; the readiness score is the one that means something on any given day.

Our presence score is flat. What moves it?

First check whether the sample is stable and large enough to interpret. Then review which sources are cited, what buyer questions lack a useful answer and where credible third-party coverage is missing. scaile manages the research, writing, fact-checking, approval and publication of that content. Publishing is an intervention to test, not a guaranteed score increase.

Sources

More articles

See where your brand stands in AI search

Book a call and we walk through your category: the questions your buyers ask, who gets cited today, and where the opportunity is.