AI visibility measurement
Designing a brand-blind AI visibility baseline
Build a repeatable AI visibility baseline with non-branded buyer questions, fixed markets and interfaces, raw evidence, separate scores and honest unavailable states.
Published by SALIENS · Published
Measure discovery without giving the answer away
If every test prompt contains your brand name, the exercise measures how a system describes an entity it has already been given. It does not show whether a buyer asking a category, comparison, evidence or fit question could discover that brand. A brand-blind baseline therefore begins with non-branded buyer questions and keeps brand-prompted accuracy tests in a separate register.
The two registers answer different questions. The discovery baseline asks whether the brand, its website or useful evidence appears when the buyer has not named it. The accuracy baseline asks whether facts are correct when the brand is explicitly named. Do not merge them into one visibility score: a brand can be accurately described but not discovered, or mentioned during discovery with inaccurate details.
Complete an eight-field baseline register
Freeze the method for a reporting window before collecting observations. The register below makes the conditions inspectable and gives a supplier a concrete acceptance format rather than a screenshot selection.
| Field | What to fix or record | Acceptance check |
|---|---|---|
| 1. Buyer decision | Audience, need, category, market and the decision the answer should support. | Every question has a stated procurement purpose, not just a keyword. |
| 2. Question set | Exact non-branded wording, question family, inclusion reason and stable question ID. | The brand, domain, product names and distinctive slogans are absent unless the task is explicitly an accuracy test. |
| 3. Market and language | Country or region, language, and any location setting that can be confirmed. | Results from different markets or languages are not pooled without a declared rule. |
| 4. System and interface | Provider, product or interface label shown at test time, access route and available version information. | The record names what was actually used; it does not infer an undisclosed model. |
| 5. Session conditions | Date and time, signed-in or signed-out state, new or continuing conversation, and any known personalisation or location condition. | Another reviewer can reproduce the stated setup, with unknown conditions marked unknown. |
| 6. Raw observation | Full question and answer, visible links or citations, screenshot where permitted, failures and unavailable states. | Evidence is retained before interpretation; a missing answer is not silently scored as no visibility. |
| 7. Classification | Brand mention, own-domain link or citation, third-party source, factual accuracy, alternatives shown and material caveats. | Each field has a written definition and can be reviewed independently. |
| 8. Review and change | Reviewer, reporting window, retest rule, question or interface changes, exceptions and approval. | Method changes are versioned; old and new conditions are not presented as a like-for-like trend. |
Build questions around real buying tasks
Use a small, balanced set that a buyer could plausibly ask before knowing a supplier. Useful families include category discovery, approach comparison, evidence and risk, organisational fit, and the next step. For example: what should a university include in an AI visibility audit; how should a UK organisation compare ongoing GEO services; or what evidence should a bilingual GEO provider supply. Treat these as templates and adapt them to the actual market and decision.
Avoid manufacturing dozens of near-identical prompts to inflate the denominator or hunting for wording that produces a preferred answer. Google's current guidance warns against creating large quantities of pages for query variations and says generative AI features can use query fan-out and varying models or techniques. A baseline should therefore document a useful sample under fixed conditions, not claim to replicate an internal ranking system.
- Category: Which providers or approaches could solve the defined need?
- Comparison: What should a buyer compare before selecting a service?
- Evidence: What proof and limitations should a credible provider publish?
- Fit: Which option suits the stated organisation, market, language or delivery constraint?
- Enquiry: What information should the buyer prepare before contacting a provider?
Report observations without inventing precision
State the denominator and preserve the raw evidence. A simple report can show how many questions produced a brand mention, an own-domain link or citation, an accurate material fact, or no usable observation. Keep those columns separate. Mark a system error, access limitation or unavailable response as unavailable, not as zero and not as a failed brand result.
Compare reporting windows only when the question set and relevant conditions remain stable. If a platform, interface, region, language, session method or scoring rule changes, record the change and start a new comparable series where necessary. NIST describes its AI Risk Management Framework as support for incorporating trustworthiness into the use and evaluation of AI systems; for procurement, the practical implication is to make context, method, evidence, review and limitations explicit.
Keep external observations separate from business outcomes
An observed mention or citation is not the same as a website visit, qualified enquiry or sale. Keep delivery evidence, AI observations, search data, website behaviour and commercial outcomes in separate fields with their own source and date. Google states that AI-feature visibility is not guaranteed even when requirements and best practices are met, and its Search Console reporting covers Google Search data rather than every external AI system.
Ask a provider to disclose the fixed question register, markets and languages, tested interfaces, session conditions, raw-evidence format, classification rules, review cadence, treatment of unavailable observations and every method change. Reject a proposal that replaces those details with a single unexplained score or guarantees inclusion, citations, rankings, traffic or enquiries.
How SALIENS can establish and use the baseline
SALIENS provides GEO audits, AI visibility monitoring, content and technical improvements, and ongoing GEO delivery for UK and international organisations, including scoped English and Simplified Chinese work. We can define a brand-blind question register, record the observation conditions and evidence, separate discovery from accuracy, and use documented gaps to prioritise controllable website work.
Email info@oxfordintelligence.co.uk with your website, target markets and languages, priority buyer decisions, known competitors or alternatives for context, and the systems you want observed. SALIENS is operated by Oxford Intelligence Limited. A baseline records observations under stated conditions; it does not guarantee crawling, indexing, mentions, citations, rankings, visits or enquiries.
Sources
- Google Search Central: AI features and your website (checked 1 October 2026)
- Google Search Central: Optimizing your website for generative AI features (checked 1 October 2026)
- NIST: AI Risk Management Framework (checked 1 October 2026)
Prepared with AI assistance using the linked sources and SALIENS service information.
