Methodology
How we measure AI visibility
A single AI visibility score — even an accurate one — hides more than it reveals. Whether AI knows about you, whether it recommends you in constrained searches, and whether it surfaces you in open-ended queries are three meaningfully different things. Blending them produces a number that is hard to interpret and impossible to act on.
We measure all three separately. Each construct has a distinct research rationale, distinct query design, and distinct interpretation. Together they give you a complete picture of where you stand and what to work on first.
Knowledge & Perception
We ask AI directly about you — your name is in the prompt. This answers: does the model know who you are, and does it describe you accurately?
Findability Rate
We send constraint-narrowed prompts where you are not named, but the answer set is small. This answers: does AI recommend you when a buyer describes your exact niche?
Discovery Rate
We send open-ended category prompts where your name never appears. This answers: does AI surface you in broad searches, without any description narrowing the field?
Framework
Why three constructs, never one number
Measurement researchers call this the problem of construct validity. When you blend different phenomena into one score, you risk measuring none of them accurately. A 40% blended score might mean you are highly findable but completely unknown — or well-known but impossible to find in the queries that matter. Those are opposite problems requiring opposite interventions.
Our three-construct design follows a layered measurement model common in brand research: awareness (Knowledge) → consideration (Findability) → organic recall (Discovery). They are ordered by how much prompting was given to elicit the mention, from fully aided to fully unaided.
These three numbers are reported in order of interpretability: Knowledge first (most controllable), then Findability (niche-buyer proxy), then Discovery (organic reach, hardest to move). We do not average them.
Construct 1 — Aided
Knowledge & Perception
What we measure: We ask AI directly about you — your name is in the prompt. This answers: does the model know who you are, and does it describe you accurately?
Why it matters: Aided measurement is the foundation of brand awareness research. Before measuring whether you are found unprompted, you need to establish whether there is anything to find. A model that knows nothing about you cannot recommend you.
How we measure it: We send queries like “Tell me about [Name] and their work in [Category],” then evaluate coverage (did the model have information?), accuracy (were the facts correct?), and framing (positive, neutral, negative, or dismissive). Coverage and framing come from a secondary LLM judge calibrated against human evaluations.
On 0%
0% Knowledge coverage is not embarrassing — it is the accurate starting point for anyone not yet well-represented in AI training data. It tells you where to focus first.
Construct 2 — Semi-aided
Findability Rate
What we measure: We send constraint-narrowed prompts where you are not named, but the answer set is small. This answers: does AI recommend you when a buyer describes your exact niche?
Why it matters: This is the most practical buying-intent signal. Real buyers describe what they need — specialty, location, audience — and AI returns a short list. If you are not in that list, you are invisible to buyers who match your profile exactly.
How we measure it: We generate prompts around your specialty, geography, and target audience using real buyer phrasing. Example: "Who are the best estate planning attorneys in Phoenix for small business owners?" Your name never appears in the prompt — only in the response if AI recommends you.
On 0%
0% Findability is the most common starting score. It is also the most improvable: targeted content that signals your specialty and location tends to move this score.
Construct 3 — Unaided
Discovery Rate
What we measure: We send open-ended category prompts where your name never appears. This answers: does AI surface you in broad searches, without any description narrowing the field?
Why it matters: Unaided recall is the gold standard for AI brand discovery — the equivalent of organic search ranking. If AI names you in open-ended queries, you are competing at the top of your category without any qualification. Most professionals score 0% here, which is honest: AI mentions only the most widely documented entities in broad category searches.
How we measure it: We generate prompts like "What are the best [category] options in [region]?" and "Who should I consider for [use case]?" across many phrasings. Mentions here represent the highest-intent AI recommendation a professional can earn.
On 0%
0% Discovery is common and expected for anyone without widespread media coverage or extensive AI-indexed citations. It is not failure — it is the baseline almost everyone starts from.
Response analysis
Not every mention counts the same way
For Discovery and Findability, we distinguish between genuine recommendations and weaker forms of presence. A casual name-drop is not the same as a clear recommendation — and treating them as equal would inflate the score.
Recommended
The assistant explicitly steers the user toward you. Clear buying signal.
Counts
Listed
You appear among real options the user could choose from.
Counts
Mentioned in passing
Your name appears but not in a way that would send business your way.
Does not count
Absent or dismissed
You are missing, or the assistant actively redirects away from you.
Does not count
Classification is done deterministically first (exact match, position heuristics), with an LLM judge used only for ambiguous cases. This approach follows standard content classification practice — fast rules handle clear cases; inference handles edge cases with a consistent rubric.
Score quality
A rate without a range is not a measurement
Every rate we report — for each construct, per model, across all models — comes with a 95% Wilson confidence interval. The Wilson score method is preferred over the naive proportional interval because it performs accurately at low sample counts and at extreme rates (near 0% or 100%), which is exactly where AI visibility scores for professionals tend to cluster early on.
A narrow interval means the rate is solid. A wide interval means you need more runs before the number is worth acting on. We show both so you can judge which is which.
Drift alerts only fire when one confidence interval stops overlapping the previous one — a statistically meaningful change, not a random fluctuation.
Question set
One question is not a measurement
Real buyers phrase the same intent many different ways. A measurement system should cover that variation, not lock in on a single phrasing. We generate a broad panel of Discovery and Findability prompts from real buyer patterns — category + location, use-case framing, audience-narrowed, comparison-style — so the result reflects the distribution of actual queries, not a lucky or unlucky phrasing.
Knowledge prompts are template-based rather than harvested — they are designed to test model coverage and framing systematically, not to sample buyer behavior.
Model coverage
We watch the models that shape buying decisions
Paid plans cover the AI assistants that currently handle the largest share of consumer queries — ChatGPT, Claude, and Gemini. Discovery and Findability rates are computed per model, so you can see whether one model is a particular blind spot.
Custom plans can extend to additional models via bring-your-own-key. The three- construct framework applies regardless of which models are included.
