What is a good AI visibility score? What the number measures and where it stops being useful

In short
Every tool in this market returns a number called an AI visibility score, and no two of them compute the same thing. This page separates the two families these scores fall into, gives you a usable reading for each, and says plainly what the number does not tell you. scaile publishes the arithmetic behind its own score for the same reason: a diagnostic you cannot check is a diagnostic you cannot act on.
Key takeaways
No standard exists. Semrush returns a 0 to 100 benchmark, Peec and Profound both return percentages but divide by different denominators, and Ahrefs Brand Radar returns no single score at all. Four vendors, four numbers, none of them comparable with another.
Scores fall into two families. Readiness scores are computed from your own site and are deterministic in what they ask: the same checks over the same evidence give the same score. Presence scores are computed from sampled model answers and move on their own.
The movement is not a rounding error. Answers are assembled live from retrieval, and AI search results are, in Semrush's own words, fast-changing and highly personalised — so a presence number sampled once is a single draw, not a level.
scaile's free Health Check is a readiness score: 103 checks, 0 to 100, a letter grade, and a stated count of how many checks actually ran. The AI Visibility Check is a presence measurement: ten buying questions across three assistants.
A readiness score is not a ranking and does not predict traffic. It measures whether an assistant can read and attribute your pages at all, which is the condition for everything that comes after.
There is no industry-standard AI visibility score, so the honest answer depends entirely on which instrument produced the number. On a readiness score, computed from your own site, 90 and above means nothing rated critical or high is failing, 70 to 89 means you are readable with named gaps worth closing, and below 70 is worth reading top-down by severity, because a single access-layer problem can account for a cluster of failures. On a presence score, computed by asking assistants real questions and counting how often you are named, there is no good number in the abstract at all — only your position against the same competitors in the same questions.
Anyone who quotes you a benchmark without naming the instrument is quoting noise.
Both numbers describe a gap; neither closes one. A readiness score improves with a week of technical work. A presence score only moves when a complete, sourced answer to each buying question exists on your own site and stays current. That production half is what scaile runs as managed infrastructure: a named AI Search Strategist, research from your own knowledge base, every fact checked, your team approving before anything publishes, live in 14 days.
Why is there no standard AI visibility score?
Because every vendor invented its own, and the definitions genuinely differ.
Semrush’s AI Visibility Score is documented as “a benchmark score (0–100) showing how often your brand appears in AI-generated answers compared to competitors.” It is not built by asking models your questions: Semrush draws on a database of over 317 million prompts and responses across ChatGPT, Gemini, AI Overviews and AI Mode, captured from clickstream data rather than through model APIs, covering 117 regional databases. The prompt database refreshes daily on a rolling basis; the brand reports built on it refresh weekly, so a dashboard you open on Tuesday is not showing you yesterday. The exact weighting is proprietary.
Peec defines its Visibility Score with an explicit formula: responses mentioning your brand, divided by total responses, times 100. Profound also reports a Visibility Score as a percentage, but divides brand appearances by the total responses containing at least one brand — a different denominator, from prompts Profound sends to answer engines daily. Run the same brand through both on the same day with the same questions and the two numbers will disagree, and both will be correct.
Ahrefs took a fourth road. Brand Radar reports brand mentions across the major AI platforms — the list has grown twice in 2026, so check the product page for what is covered today — plus an AI Share of Voice against competitors, with no 0 to 100 score at all. Its own guidance says AI platforms are non-deterministic and that you should track the range of your share of voice rather than a point value.
Two things follow. First, “our AI visibility score is 62” is not a fact about your company until you say whose 62 it is. Second, no industry benchmark exists to compare it against, because a benchmark requires an instrument everyone shares. Semrush says so itself in its own documentation: AI search responses are fast-changing and highly personalised, “which means no platform can provide exact numbers.”
Even the research had to invent its own yardstick. The GEO paper that established this field measured visibility with three metrics it defined from scratch — a normalised word count, a position-adjusted word count that decays exponentially with where the citation appears, and a subjective impression score scored by a language model — precisely because a ranking position, the metric classic search runs on, does not describe a generative answer.
What are the two families of AI visibility score?
Every score on the market is one of two things, and confusing them is the most expensive mistake in this whole topic.
Readiness scores are computed from your own site. Can the AI crawlers reach you? Can a machine parse an answer out of the page? Does the page carry a date, an author and a source, so the answer can attribute it? These are questions about files that exist on your server right now. A readiness score is deterministic in what it asks, not in what it can reach: the same checks over the same evidence give the same score, and the ran-count the report prints is what tells you whether the evidence was the same. A third-party API that times out, a page that would not render, a robots.txt that came back a 503 — each one takes a check out of the denominator, and a score built on 62 checks is not the same measurement as one built on 95.
Presence scores are computed from sampled model answers. Ask an assistant a question, read the answer, record whether you were named. Repeat. These measure something readiness cannot: whether you actually come up when a buyer asks. They also move on their own, and more than most dashboards admit.
Why does a presence score move on its own?
Because the answer is assembled live, and almost nothing in that assembly is held fixed between two runs.
The dominant cause is retrieval, not the model. Semrush states it plainly in its own documentation: AI search responses are fast-changing and highly personalised, which is why “no platform can provide exact numbers.” Ahrefs gives the same warning from the other side, telling Brand Radar users that AI platforms are non-deterministic and that the honest reading is a range rather than a point value. Between your Monday sample and your Thursday sample, the search index behind the assistant has churned, the pages it retrieved have changed, and the personalisation layer has treated the two requests differently.
Underneath that sits a smaller, more surprising floor. Thinking Machines sampled the same prompt 1,000 times at temperature zero — the setting that is supposed to make output deterministic — on an open-weights model and got 80 distinct completions, the most common occurring 78 times, with the first divergence at the 103rd token. The cause was not randomness in the sampler but batch-size variation on the serving hardware, which changes numerical results depending on how many other people happen to be querying at the same moment. The important half of that result is the fix: with batch-invariant kernels, all 1,000 completions came back identical. So this is not evidence of irreducible randomness in language models. It is evidence that even the floor case is not reproducible on ordinary serving infrastructure, before retrieval and personalisation add their much larger movement on top.
The GEO researchers handled the same problem the honest way: they sampled five responses per query at temperature 0.7 “to reduce statistical deviations,” on GEO-bench, a benchmark of 10,000 queries. That is what a presence measurement costs if you want it to hold still.
So when a dashboard shows your presence score moving from 34 percent to 31 percent week on week, the first question is not what caused the drop. It is whether three points is even outside the instrument’s own noise floor. On ten questions asked once each, one answer flipping moves the number ten points. Nothing about your website changed.
This is why scaile’s AI Visibility Check hands you the underlying answers rather than only a headline number: what the assistants actually said is checkable, and a percentage derived from a single pass is not. Continuous tracking with enough repetition to separate signal from noise is a different job, and it is what the reporting side of the platform does.
What counts as a good readiness score?
This is the family where a number in the abstract does mean something, because the same checks over the same evidence always produce the same score.
On the 0 to 100 scale the free AI Search Health Check uses:
90 and above is an A. Nothing rated critical or high is failing. Whatever is still open is cosmetic or optional, and whatever is missing from AI answers is missing for content reasons rather than technical ones.
70 to 89 is a B or a C. You are readable, with named gaps worth closing. Typically a mix: schema on some page types and not others, a sitemap that misses a section, dates present but authorship absent. Each fix is small and none of them is blocking.
Below 70 is where triage beats a to-do list. Read the result top-down by severity rather than working down the list, because a single problem in the access layer can account for a whole cluster of failures underneath it.
The distinction that matters at that layer is which bot you blocked. Blocking the retrieval agents — OAI-SearchBot, PerplexityBot — is the costly block: those are the agents that fetch a page at answer time, and in scaile’s own engine a robots.txt that disallows the retrieval crawlers together trips a score gate that caps the result outright, no matter what else passes. Blocking the training crawlers — GPTBot, ClaudeBot, CCBot — is a rights decision about your content being used to train future models. It is a legitimate one to make, and it barely moves whether ChatGPT or Perplexity cite you today, because citation runs through the retrieval agents. Fix the gate and the cap lifts; fix a training-crawler line and you have changed your licensing posture, not your visibility.
A score is only trustworthy if it shows its arithmetic. The Health Check reports how many checks passed out of how many actually ran, not out of 103, because the free pass reads your homepage, your robots.txt and your sitemap — enough to settle around 60 of the checks — and the rest abstain rather than guess. A check that cannot reach a verdict should lower your confidence in the score, not silently lower the score.
What does the score not tell you?
A readiness score is not a ranking, and it does not predict traffic. It is a precondition test.
There is no external way to validate a correlation between a readiness score and AI-driven visits, because the data does not exist publicly. Google states that pages appearing in AI Overviews and AI Mode are folded into overall Search traffic in Search Console, with no separate AI-features metric. And the most valuable outcome leaves no trace at all: a brand named in an answer that the reader never clicks produces no analytics event anywhere, while referrer handling on the clicks that do happen differs from assistant to assistant. Any vendor claiming their score predicts revenue is claiming something nobody in this market can currently measure.
What the readiness score does tell you is whether the condition for everything else holds. Google’s own requirement for appearing in AI Overviews is exactly one thing: the page must be indexed and eligible to be shown with a snippet. Google adds explicitly that no new machine-readable files, no AI text files and no special schema are needed. That is a low bar, and sites do fail it. When all three of the access checks fail together, the cause usually sits at the network edge rather than in the site: a CDN or WAF rule blocking a crawler that robots.txt permits.
Score 95 and you can still be absent from every answer in your category, because readiness measures whether you could be cited, not whether the answer to your buyers’ question has been written. That gap is a content problem, and it is the one that takes actual work: dozens of concrete buying questions, each with a complete, sourced, current answer in your own company’s language. Finding which questions those are is what customer-intent research is for, and the causes of invisibility sit in a predictable order once the technical layer is clean.
How should you use the two numbers together?
Run the readiness score first and treat it as a gate. It takes a minute, and it tells you whether a presence measurement would even be interpretable. Measuring how often ChatGPT names you while OAI-SearchBot is blocked tells you nothing you did not already know.
Then measure presence, and read it as a position rather than a level. Named in most answers and named early is strong. Named late, or only when the question is narrow, is a real gap. Never named in your own category is the finding that matters, and at that point the readiness score tells you which of the two causes you are looking at.
Both scaile checks are free and neither needs an account. The Health Check scores readiness on the page. The AI Visibility Check writes the ten buying questions and shows them on the page for free; the answers themselves, the competitor leaderboard and the sources the assistants read come back in the emailed report, which needs an email address but no signup. If you want the structural view first, the Sitemap Visualiser draws what the crawlers can and cannot reach before any scoring happens.
FAQ
Is 70 a good AI visibility score?
On a readiness scale it is a passing grade with named gaps: readable, nothing catastrophic, several small fixes outstanding. On a presence scale, 70 percent of answers naming you would be exceptionally strong — which is exactly why the two numbers must never be compared. Always ask which instrument produced the number.
Can I compare my Semrush AI visibility score with another tool’s?
No. Semrush computes a 0 to 100 benchmark from a clickstream prompt database, Peec divides brand mentions by all responses, and Profound divides by responses containing at least one brand. Different inputs, different denominators, different numbers for the same brand on the same day.
Why does my AI visibility score change when nothing changed on my site?
Because presence scores sample live answers, and those answers are rebuilt from live retrieval on every run, over an index that churns, with personalisation on top. Semrush and Ahrefs both say so in their own documentation. A readiness score computed from your own files does not move unless your files, or what the checker could reach, do.
Does a high AI visibility score mean more traffic?
There is no public evidence for that, and there cannot be yet: Google does not break out AI-features traffic in Search Console, and a mention that is never clicked produces no analytics event at all. Treat a readiness score as a precondition test, not a forecast.
How is scaile’s AI visibility score calculated?
103 checks across technical foundation, content architecture, AI optimisation, meta and social, and experience and trust. Each passes or fails on its own, they roll up into a 0 to 100 score with a letter grade, and the result states how many checks actually reached a verdict rather than implying all 103 did.
Should I track an AI visibility score every week?
Only if the measurement repeats enough times to separate a real move from noise. A weekly single pass over ten questions is a coin-flip dressed as a trend line; the readiness score is the one that means something on any given day.
Our presence score is flat. What moves it?
Published answers, at a cadence. A presence score only rises when the questions your buyers ask are answered completely and with sources on your own site, and when those answers stay current. scaile runs that as managed infrastructure: a named AI Search Strategist, research from your knowledge base, fact-checking on every claim and your team approving before publication, live in 14 days. Building Radar doubled its qualified inbound leads in 90 days on it.
Sources
- Semrush: AI Visibility Metrics. Definition of the AI Visibility Score as a 0 to 100 competitive benchmark.
- Semrush: where the AI Visibility Toolkit data comes from. 317 million prompts and responses, clickstream rather than APIs, 117 regional databases, the daily rolling prompt refresh against weekly brand reports, and the statement that no platform can provide exact numbers.
- Peec AI: Visibility metric documentation. Visibility Score as responses mentioning the brand divided by total responses.
- Profound: Answer Engine Insights overview. Visibility Score divided by responses containing at least one brand; prompts sent to answer engines daily.
- Ahrefs Brand Radar and the Brand Radar announcement. The current platform list, AI Share of Voice, no 0 to 100 score, and the guidance to track a range because AI platforms are non-deterministic.
- Defeating Nondeterminism in LLM Inference, Thinking Machines. 1,000 samples at temperature zero, 80 distinct completions, most common 78 times, first divergence at token 103, batch-size variance as the cause, and all 1,000 completions identical once batch-invariant kernels are used.
- GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024. The three purpose-built visibility metrics, GEO-bench at 10,000 queries, and five samples per query at temperature 0.7 to reduce statistical deviation.
- Google Search Central: AI features and your website. Indexed and snippet-eligible as the only requirement, no special files or schema, and AI-features traffic reported inside overall Search traffic.
- OpenAI: crawler overview. OAI-SearchBot, GPTBot, ChatGPT-User and OAI-AdsBot, and what each one is used for.



