← All insights

playbook · AI Visibility

How to test AI visibility without ranking theater

A useful AI visibility review preserves the question, answer, cited source, surface, date, and commercial relevance.

A growing number of businesses now ask some version of the same question: when a qualified buyer asks ChatGPT, Claude, Gemini, Perplexity, or Google's AI results who to trust, do we appear, and are we described accurately? It is the right question. The problem is how it usually gets answered.

The common approach is ranking theater: someone types two or three flattering prompts, screenshots a good answer, and declares visibility. Or a vendor produces a proprietary score with no visible inputs, no dates, and no way to reproduce the result. Both feel like measurement. Neither survives contact with a serious buyer, and neither tells you what to fix.

A useful AI visibility review has a simpler and stricter shape. It preserves five things for every observation: the question asked, the answer given, the sources cited, the surface it was asked on, and the date. Everything else is derived from those five.

Start with the questions buyers actually ask

Do not begin with your brand name. Begin with the decisions your buyers are making. A medspa's buyers ask about a specific treatment in a specific city, whether it is safe for their situation, and how to choose a reputable provider. A law firm's referral sources ask who handles a narrow matter type well. Brand-name prompts mostly confirm that the engine can read your website. Decision prompts reveal whether you are part of the answer when it matters commercially.

Keep each question specific to a service, market, specialty, or decision. Twelve well-chosen questions beat one hundred generic ones, because you will actually re-run twelve on a schedule.

Record answers you can re-review, not impressions

For each question, capture the answer with enough fidelity that someone else could review its meaning later, and preserve every cited source with the observation date. A mention without context is not qualified visibility. "Mentioned third, described as a general practice, citing an outdated directory" is an observation you can act on. "We showed up" is not.

Dates matter more than most teams expect. AI answers change as models, retrieval systems, and the underlying web change. An observation without a date cannot be compared honestly against a later one, and comparison is the entire point.

Grade what matters

With the raw observations preserved, grade each one on a small set of dimensions:

  • Presence: does the business appear at all for this question?
  • Accuracy: are the services, locations, and claims described correctly?
  • Source quality: which sources does the answer rely on, and do you control or influence any of them?
  • Competitive ownership: who is being recommended instead, and on what cited basis?
  • Usefulness: does the answer give the buyer a clear next step that includes you?

The source dimension is where the work usually comes from. If the engines consistently cite a directory with stale information, that directory is now a priority. If they cite competitors' educational pages for questions in your specialty, that is a content gap with a commercial address attached.

Repeat the same set on a cadence

A single review is a snapshot. The value compounds when the same question set is re-run on a defined cadence, monthly for most local and professional-services businesses, and the grades are compared period over period. Changes can then be tied, cautiously, to the work performed in between: structured data shipped, directory corrections, new evidence pages, media placements.

Cautiously is the operative word. Engines change for their own reasons. An honest program reports movement alongside what was done, without claiming sole causation, and flags observations that were probably influenced by outside events.

What to refuse

Refuse scores that cannot show their inputs. Refuse guarantees of AI citations or rankings; no ethical provider controls what an engine says. Refuse reviews that only test flattering prompts. And refuse conclusions drawn from a single day's observations. The variance across days and sessions is real, and pretending otherwise is theater with extra steps.

What remains after all that refusal is smaller, slower, and dramatically more useful: a dated, reproducible record of how the engines describe you, what they rely on when they do, and a prioritized list of the sources and pages worth fixing next. That is a measurement program a serious operator can defend, to a partner, to a board, or to themselves.

Apply this to your operating system.

Use the relevant five-minute assessment to identify the first evidence-backed improvement.

Discuss the use case
Book a strategy call