Standards
How AIRR maps to the IAB's AI visibility framework
In August 2026 the Interactive Advertising Bureau published “Measuring Visibility in the AI Era”, a framework meant to give the AI-visibility measurement market a shared vocabulary: the “4 P's” (Presence, Prominence, Portrayal, Persuasion) and a distinction between directional data — good for early signal and internal briefings — and decision-grade data, rigorous enough to base budget and strategy on.
We built our methodology before this framework existed, so this page is our own mapping of what we already do onto IAB's vocabulary — not a certification issued by IAB, and not a claim of formal compliance. Where we fall short of a P or a criterion, we say so below rather than stretch the mapping to fit.
The 4 P's
Where each P lands in AIRR today
Presence
IAB: “Do you appear in an AI response at all?”
Our Discovery Rate (unaided) and Findability Rate (semi-aided) are both presence measures at different levels of prompting, and Knowledge coverage answers whether the model has any information about you at all — the precondition for presence anywhere else.
We report these as three separate rates rather than one "presence" number, because whether you're absent from an open-ended category search and whether you're absent from a query naming your exact niche are different problems with different fixes.
Prominence
IAB: “Where, and how substantively, do you appear?”
Every mention is classified by response role — Recommended, Listed, Mentioned in passing, or Absent/dismissed — and only the first two count toward a rate. A name-drop doesn't inflate your score the way it would under a simple mention count.
At the source level, Source Selection Diagnosis goes further: when AI cites a competitor's page instead of yours, we diagnose which stage of the funnel — access, coverage, vocabulary, semantic relevance, citability, or freshness — caused the drop, and cite the specific evidence for each stage.
Portrayal
IAB: “In what context, and with what accuracy?”
This is what our Knowledge construct measures directly: coverage (does the model have information about you), accuracy (are the facts correct), and framing (positive, neutral, negative, or dismissive) — scored by an LLM judge calibrated against human evaluations.
We don't currently publish a standalone hallucination-rate metric distinct from the accuracy score above; today accuracy and framing are captured together per response rather than decomposed into IAB's finer-grained sub-metrics.
Persuasion
IAB: “Does visibility drive action back to you?”
This is a real gap. AIRR measures what AI says and cites, not what happens after — we don't currently track AI-referred site traffic or crawler activity from AI agents against your own analytics or server logs.
It's on our roadmap, not shipped. We'd rather say that plainly than pad the framework with a metric we can't back with data yet.
Data quality
What makes a rate decision-grade, not directional
IAB's decision-grade bar isn't a single number — it's a checklist of how a measurement was produced: sample size, query volume, prompt-type coverage, testing cadence, reproducibility, and data validation. Here's how each one shows up in AIRR specifically.
Sample size & query volume
Discovery and Findability rates are generated from a broad panel of prompts per construct — not a single phrasing — and every rate we report carries a 95% Wilson confidence interval, chosen because it stays accurate at the low sample counts and extreme rates (near 0% or 100%) where AI visibility scores tend to sit.
Prompt-type coverage
Coverage is designed in, not sampled after the fact: aided (Knowledge), semi-aided (Findability), and unaided (Discovery) prompts each use a distinct query design built from real buyer phrasing patterns — category + location, use-case framing, audience-narrowed, comparison-style — rather than one prompt style stretched to cover every use case.
Testing cadence
Measurement runs in recurring windows, not a one-off snapshot, so a score is a point on a trend rather than a single unrepeatable draw.
Reproducibility
Response classification is deterministic first — exact match and position heuristics — with an LLM judge invoked only for ambiguous cases and calibrated against human evaluation. That keeps run-to-run variance from being judge noise rather than real movement.
Data validation
Drift alerts only fire when two windows' confidence intervals stop overlapping — a statistically meaningful change, not a random fluctuation. And every Source Selection Diagnosis recommendation is tagged with its own evidence tier (Observed fact, Strong inference, Heuristic, or Speculative), so a claim's confidence is never hidden inside a confident-sounding sentence.
Why publish this
Ask any AI-visibility vendor these questions
IAB built this framework because the market couldn't compare vendors on anything but marketing claims — one survey cited alongside the framework's release found only 16% of brands currently track AI visibility at all, in large part because it's been unclear what a given vendor's number actually measures or how reliable it is.
Our answer is to publish the mapping above and our full methodology rather than a single opaque score. If you're evaluating us against another vendor, ask both of us for this same breakdown — sample sizes, prompt design, testing cadence, reproducibility, and what's honestly still a gap. That comparison is the point of the framework existing at all.
