The Proof Standard for AI Visibility Claims
By Firoz Azees
One of the 10 Dubai agencies we measured shows any AI-answer evidence on its service page. The category has no proof standard, so here is one: six requirements, all cheap to meet, and a worked example carrying our own failing score.
5 min readA category without a proof standard runs on assertion. We measured 10 Dubai agencies across 8 buyer questions and found exactly 1 showing any measured AI-answer evidence on its own service page. The other 9 sell AI visibility on claims. This page proposes the standard, which costs nothing to meet and is the reason we can publish our own failures.
The short version
Six requirements make an AI-visibility claim checkable: a date, a named engine and locale, the exact questions asked, the full saved answers, named and cited counted separately, and the vendor's own score included. Every one is free. Across the 31 GEO products and 10 agencies we track, the category norm is a blended share-of-answer figure with none of the six, which is why a 300% improvement claim carries no information: a 300% lift on 1 mention is 4 mentions.
The six requirements
1. A date
Generated answers are assembled at query time from a changing pool. An undated claim is not a measurement. Our own readings carry three dates: 21 July, 7 September and 12 September 2026.
2. A named engine and locale
Answers differ by engine and by where you ask from. "AI search" is not a measurable object; Google AI Mode from a UAE browser is.
3. The exact questions
Publish them so the buyer can retype them. A question set held back as proprietary is protecting the claim, not the method.
4. The full saved answers
Screenshots crop. A stored answer with its source list is the artefact that survives disagreement.
5. Named and cited counted separately
AI Search Visibility is not one number.
Named and cited are different outcomes with different causes. Collapsing them hides the most useful diagnostic a buyer has, as our own panel showed when one agency was cited in 50% of answers while named in 12.5%.
6. The vendor's own score
A referee who hides their score is not a referee.
The standard as a table
| Requirement | Cost to meet | Met by the category |
|---|---|---|
| Date on every reading | Zero | Rare |
| Named engine and locale | Zero | Rare |
| Published question set | Zero | Almost never |
| Stored full answers | Near zero | Almost never |
| Named vs cited split | Zero | Almost never |
| Vendor's own score | Zero, and uncomfortable | 1 of 10 shows any evidence |
A worked example, carrying our own failure
Here is the standard applied to us. Engine: Google AI Mode. Locale: UAE. Dates: 21 July, 7 September, 12 September 2026. Questions: the 8 buyer questions listed in full on our Dubai agency measurement. Result: Ivanooo named 0%, cited 0%, on all three dates. Wider panel: Share of Recommendation 0.0 across 62 queries (16-18 September 2026), with 59 of 62 returning us absent, 2 citing us and 1 naming us.
That is a failing score published to the same standard we would use for a win. As Firoz Azees puts it: "a measured zero is a work order; a flattering estimate is a story you tell yourself."
How to apply it in a sales conversation
- Send the six requirements before the call. Vendors who can meet them will; the rest will reframe.
- Ask for two dated readings, not one. In a category where leaders turned over 100% inside 52 days, a single date proves nothing.
- Ask which questions they expect to lose. A practitioner can name them. A salesperson cannot.
- Ask for their own score. The answer tells you more than the deck.
- Keep your own baseline. The conversation changes when you hold the measurement rather than receiving it.
Questions buyers ask
Is this standard too demanding for a small agency? Every requirement is free. A two-person shop can meet all six this afternoon; the barrier is willingness, not resources.
What about proprietary methodology? Interpretation can be proprietary. The questions and the dates cannot be, because without them no claim is checkable.
Do the software tools meet it? Not in the category norm. Of the 31 GEO products we track, blended share-of-answer reporting is standard, which fails requirements 3, 5 and 6 at once.
Why does the vendor's own score matter to me? Because it is the only evidence that the method survives contact with a hard case. Anyone can show a flattering client.
Does meeting the standard guarantee results? No, and the limits of this discipline are worth reading alongside it. The standard guarantees that you can tell whether results happened.
How does this relate to the research? The Princeton GEO study measured passage-level effects up to 40% under controlled conditions. A controlled benchmark is the closest published analogue to what this standard asks vendors to produce in the field.
What is the fastest way to start? Take one reading yourself, to the six requirements, before any vendor conversation. The method is in how to verify an agency's claims.
At Ivanooo, Firoz Azees runs Distinctiveness Engineering for the AI-answer era: measuring who the engines name, cite and recommend, then engineering the gap between listed and chosen. Our pricing is published and our score is above. Start with a free AI visibility check.