Run Your Own AI Visibility Test in One Hour
By Firoz Azees
The method behind every number we publish, handed over in full. Eight questions, one hour, no tools, and a baseline nobody can sell you a story about afterwards.
5 min readWe ran this method on our own market three times and published every result, including the failures. It is simple enough to hand over, so here it is in full. One hour, no software, and at the end you hold a dated baseline. That baseline is the thing that changes the next sales conversation you have, because a buyer holding a measurement cannot be sold a story about one.
The short version
Write the 8 questions your buyers type. Ask each one in a clean browser in your buyers' locale. Save the full answer and its source list with today's date. Count three things per question: are you named in the text, is your domain in the sources, and are the sources mostly platforms or mostly websites. Repeat in 60 days. That is the entire method, and it is the same one that produced our own failing score of 0% across three dates.
Step 1: write the questions
We wrote ours in the words a customer types, never category jargon. "Best AEO agency in Dubai" is a buyer question. "Generative engine optimization solutions" is a vendor's phrase. Eight is enough to be informative and small enough to finish. Ours are published in full on the Dubai agency measurement if you want the shape.
Step 2: ask them in a clean session
We ask from a fresh browser session, signed out, in the locale our buyers sit in. Personalisation and location both change answers, and a reading taken from your own logged-in account measures your history rather than the market.
Step 3: save the receipts
We copy the full answer text and the complete source list into a dated file. Screenshots crop and lose the source list, which is the half that carries the diagnosis.
Step 4: count three things
| What to count | How | What it tells you |
|---|---|---|
| Named | Your brand appears in the answer text | The engine endorses you |
| Cited | Your domain appears in the sources | The engine uses you |
| Source class | Platform vs website, per source | Whether on-domain work can win it |
Step 5: read the diagnosis
- Named and cited both zero, engine describes you wrongly — an Entity problem. Fix identity before anything else.
- Described correctly, still not named — a distinctiveness problem, the slowest and most valuable to fix.
- Named but never cited — your pages are not retrievable at passage level.
- Cited but never named — you supply the words and a rival takes the credit.
- Sources are mostly platforms — the answer is off-domain, and no page you publish reaches it.
Step 6: repeat in 60 days
This is the step that separates a measurement from a screenshot. Positions here churn: our own two-date comparison found a 100% turnover among the agencies the engine favoured, inside 52 days, with nothing visibly changed about any of them. One reading cannot distinguish a result from a re-sample.
What a good baseline is worth
It costs an hour and it changes three conversations. It tells your team which questions to attack rather than guessing. It bounds what any vendor can promise you, because you already know your starting number. And it makes every later claim checkable, which is the whole content of our proposed proof standard.
Questions buyers ask
Do I need a tool for this? No. Tools track more prompts than you can by hand, which matters for broad categories. For a first baseline, an hour and a text file are sufficient.
Which engine should I use? The one your buyers use. We measure Google AI Mode because our market sits there; the method transfers, the numbers do not.
How many questions is enough? Eight is a working minimum. Fewer and one odd answer distorts everything; many more and you will not repeat it in 60 days, which is the step that matters.
What if the answer changes when I ask twice? That is the mechanism, not an error. Generated answers are assembled per query. Record what you get and rely on the repeat reading rather than on any single run.
Should I include competitors in the count? Yes. Their named and cited counts define what winning means in your category, and the gap between their two numbers is usually more instructive than your own. In our July panel one agency held a 50% citation share against a 12.5% named share, which told us more than any single score could.
What if my score is zero? Then you have a work order rather than a verdict. Ours has been zero across three readings, published each time. As Firoz Azees puts it: "a measured zero is a work order; a flattering estimate is a story you tell yourself."
Does the research support this being worth the effort? The Princeton GEO study, Aggarwal et al., measured passage-level changes moving visibility by up to 40% under controlled conditions. You cannot tell whether such a change helped you without a before and an after, which is what this hour produces.
At Ivanooo, Firoz Azees runs Distinctiveness Engineering for the AI-answer era: measuring who the engines name, cite and recommend, then engineering the gap between listed and chosen. If you would rather not run it by hand, start with a free AI visibility check.