How to Verify an AI-Visibility Agency's Claims in 20 Minutes

By

Of 10 agencies on our Dubai panel, 1 shows measured AI-answer evidence and 0 publish pricing. Here is the diligence checklist that separates a method from a claim, and the four questions no agency selling on promises can answer.

6 min read

Every agency selling AI visibility shows you a dashboard. Almost none shows you a method. The difference matters because a dashboard supplied by the vendor being graded is not evidence, and in a category this young the gap between a real practice and a convincing deck is invisible from the outside. This is the diligence method we would run on ourselves, and it takes about 20 minutes.

The short version

We measured 10 Dubai agencies across 8 buyer questions. Exactly 1 shows any measured AI-answer evidence on its own service page and 0 publish pricing, which means 90% of the field asks you to buy AI visibility on assertion. Four questions expose that quickly: what were your numbers 60 days ago, which of my questions can you not win, which engine did you measure, and can I reproduce your result myself. Any agency that answers all four is describing a method. Most answer none.

Why the usual proof is not proof

A dashboard is not independent

The vendor chooses the metric, the sample and the window. Of the 31 GEO products we track, the category norm is a blended share-of-answer figure, which is exactly the number a vendor can move without moving your business.

A case study without a method is an anecdote

"We increased AI visibility 300%" carries no information without the baseline, the question set, the engine and the dates. A 300% lift on 1 mention is 4 mentions. The primary research on what moves generated-answer visibility, Aggarwal et al. at Princeton, measured effects of up to 40% from specific passage changes under controlled conditions, which is the order of magnitude an honest claim sits in.

An award is a purchase decision by someone else

Several Dubai agencies lead with award shelves. Awards measure submissions, not answers.

The four questions

1. What were my numbers 60 days ago?

The single most revealing question. Positions in this category churn: across our own two probe dates, 52 days apart, the leaders turned over completely, with one agency falling from a 50% citation share to 0% and another from 37.5% named to 0%. Only 20% of the July field appeared at all in September. An agency with a method has a prior reading. An agency without one has a screenshot from this morning.

2. Which of my questions can you not win?

A real practitioner can name the answers they cannot reach, usually because those answers are sourced off-domain. On our 62-query panel (16-18 September 2026), 27 of the 56 answers that cited anything cited YouTube, LinkedIn, Reddit or Quora. On-domain work cannot win those, and an agency claiming otherwise has not looked.

3. Which engine, on what date, from where?

AI answers vary by engine, by locale and by day. A claim without an engine name and a date is not reproducible, and reproducibility is the only thing separating measurement from marketing.

4. Can I reproduce this myself?

The strongest proof standard costs nothing: publish the questions so the buyer can ask them. If a vendor's evidence cannot survive you typing the same question into the same engine, it was never evidence.

The checklist

What to demand Why it matters Pass condition
Two dated measurements Position churns; one reading proves nothing Readings 60+ days apart
Named engine and locale Answers differ by both Stated explicitly
The exact question set Lets you reproduce the claim Published, not summarised
Saved full answers Screenshots crop; receipts do not Stored and shareable
Named vs cited, split Different outcomes, different work Reported separately
Off-domain share Bounds what on-domain work can win Quantified as a %
Published pricing 0 of 10 Dubai agencies do this Public rate card
Their own numbers A referee who hides their score Self-measurement shown

Run the test yourself first

  1. Write your 8 buyer questions in the words a customer would use, not in category jargon.
  2. Ask each one in a clean browser session, in your buyers' locale.
  3. Save the full answer and its source list as a dated file. That file is your baseline and your leverage.
  4. Count named and cited separately for yourself and for every competitor that appears.
  5. Take the file into the sales call. The conversation changes when the buyer holds the measurement.

What we score on our own checklist

We publish our own reading because a referee who hides their score is not a referee. Across those 8 Dubai questions Ivanooo scored 0% named and 0% cited on 21 July, 7 September and 12 September. Across the broader 62-query panel our Share of Recommendation reads 0.0, with 59 of 62 questions returning us absent. Our pricing is published. As Firoz Azees puts it: "a measured zero is a work order; a flattering estimate is a story you tell yourself."

Questions buyers ask

Is it unreasonable to demand two dated measurements from a young agency? No. A new agency can still measure your baseline today and again in 60 days before asking for a long commitment. What is unreasonable is a guarantee with no prior reading behind it.

What if the agency says the method is proprietary? The question set and the dates are not proprietary; the interpretation is. Any vendor treating "which questions did you ask" as a trade secret is protecting the claim, not the method.

How do I judge AI Search Visibility claims against SEO claims? SEO claims can be checked against rank trackers that anyone can buy. AI answers have no equivalent public record, which raises the burden of proof on the vendor rather than lowering it.

Does published pricing really matter? It is the cheapest available signal of confidence. In our Dubai panel, 0 of 10 agencies publish it, so every budget conversation in that market starts blind.

Should I hire the agency the AI already recommends? Not automatically. Being named is evidence the engine knows them, not evidence they can move your numbers. Ask how they got there and whether the method transfers.

What does good diligence amount to? An agency that shows you two dated readings, names the answers it cannot win, and hands you the questions so you can check its work yourself.

Where do I start if I have no baseline? Run the 5 steps above, then read what this discipline cannot do before setting expectations, and use our own Dubai measurement as a worked example of the format.

At Ivanooo, Firoz Azees runs Distinctiveness Engineering for the AI-answer era: measuring who the engines name, cite and recommend, then engineering the gap between listed and chosen. Start with a free AI visibility check and hold your own baseline before anyone sells you one.