Transparent by design

A score should be
explainable.

This page explains what was asked, which surfaces were measured, how recommendation states are classified, how competitors are detected, how your share of the shortlist is calculated, how prompt versions are frozen, what the system can and cannot attribute, and why VISIBLE never guarantees third-party AI rankings.

No credit card for the free check. $99 Intelligence includes 30 days of Autopilot Early Access.

The score is four measured questions, weighted and averaged.

Each part is computed per engine and averaged across the engines that measured it. A part with no evidence behind it is dropped and the remaining weights are renormalised — never rendered as a zero, because “no engine cited anything” and “nothing supported you” are different findings.

  1. 35%Are you included in the answer at all?
  2. 30%Are you actually recommended, not just named?
  3. 20%Are credible sources supporting your presence?
  4. 15%Does AI understand your brand correctly?

The headline is an integer. One pass per question per engine does not support a tenth of a point, and printing 45.8 would claim a precision the sample cannot carry.

The report and dashboard you receive today still label these four parts with the audit engine's original names. Moving that vocabulary onto the five metrics this site uses is a separate piece of work, and until it lands the two will not match word for word.

Nine modules, and what each one does.

Where the product does something, this says what it does. Where it does not do it yet, this says that instead.

Money Prompt selection and versioning

Each category has a prompt library held as a versioned file, and every question carries the version of the set it came from.

The paid audit fires thirty questions across discovery, problem and use case, comparison, trust, commercial intent, brand perception and long-tail; the free check fires eight. Ten of the thirty are the commercial core and are asked more than once.

A library that is malformed, incomplete or has two questions sharing an id is refused outright rather than loaded partially — a silently truncated instrument produces a silently wrong score.

Recommendation classifier

Every brand in every answer is graded on a five-point scale: not present, mentioned, considered, recommended, top recommendation — the ladder this site writes as Absent, Mentioned, Considered, Recommended and First Choice.

The grade is derived from the extraction layer's own reading of the answer, so it costs no extra model call and audits already run can be re-graded from stored rows.

A question counts as won from Recommended upward. Being listed among the options is not being recommended, and the gap between those two is where the metric earns its keep.

Share of Shortlist calculation

Of every recommendation handed out across the tracked brands, this is the share that went to yours — recommendations only, never raw mentions, so a brand named dismissively in every answer does not score well.

The denominator is closed over you plus your tracked competitors, which means the figure is comparable only across runs of the same competitive set. Adding a competitor later changes every historical figure, so the set is fixed deliberately.

When no brand in the set was recommended anywhere, the service reports that there was nothing to divide rather than printing 0% for everybody and implying a measured tie.

Brand Truth comparison

planned

Claims the models made about your company are graded against the text of your own website: supported, contradicted, or not addressed.

A claim your site never speaks to stays unverified. “Your site does not say” is not “the model lied”, and conflating the two is how a register of 300 harmless paraphrases becomes an accusation.

If verification cannot run — site unreachable, judge unavailable — the register is left untouched. It never invents a verdict to fill a row.

Not built yetMissing, outdated and inconsistent as distinct machine-assigned states. Today the service separates verified from wrong from unverified.

Machine Readability checks

The service crawls up to forty pages of the site and checks whether the answer engines' own crawlers are allowed in — GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest are checked by name.

It compares the raw HTML against the rendered page, reads the page structure, and looks for the brand in the third-party sources models lean on when describing a company. Every finding carries a severity and a plain explanation of what to do about it.

This is reported beside the score and deliberately kept out of it: a well-built website should not be able to disguise a brand no model ever recommends.

Action risk policies

planned

Green, Amber and Red describe how much an action is allowed to do on its own: publish automatically, wait for one-click approval, or never auto-publish at all.

The audit service that runs today measures and diagnoses. It produces findings, ranked priorities and a 90-day plan; it does not publish anything to your site.

Not built yetThe execution layer that carries these policies. Until it ships, every action VISIBLE proposes is something a person applies.

Verification windows

planned

Verification means re-asking the same frozen questions rather than asking new ones. Because the instrument is versioned, a later run is comparable to an earlier one by construction.

The recurring instrument asks its core commercial questions more than once per run. Repeats are what make “the same answer, asked again” measurable at all, and what narrows the interval around a figure.

The free check runs one pass per question per engine, so it carries no confidence interval and no trend — and says so on its own report.

Not built yetRe-test windows scheduled from an action's date, so a specific change can be measured against the questions it targeted.

Experiment outcome states

planned

Improved, No movement, Inconclusive and Regressed are the only four verdicts an experiment can end on, and three of them are not wins.

Nothing in the service assigns these states today. They are the contract for the verification loop, published here before it exists so that it can be held to it.

Not built yetThe experiment register that records a baseline, an action date and a re-test.

Known limitations and surface availability

The paid audit measures seven surfaces: ChatGPT, Claude and Gemini each answering both from memory and with live web search, plus Perplexity. The free check measures four of them. Google AI Overviews is built but switched off until a paid search provider is configured, and the report always states which surfaces actually answered.

Every engine is weighted equally. Weighting one surface above another is a factual claim about where buyers ask questions, and there is no validated data behind such a claim, so the service does not make one.

AI answers are probabilistic and third-party. VISIBLE reports measured movement and confidence; it does not guarantee causality, citations or ranking outcomes.

Discover for Free

Your customers are already asking AI.

Find out whether AI is putting you — or your competitors — on the shortlist.

  • No card
  • Report emailed to you
  • Delete your data anytime

Already know you have a problem?