voices quoted
Quoted in this piece
How to Measure AI Visibility: Count a Rate, Not a Position
AI answers have no positions. The unit that survives a re-run is a rate — appearances over runs — and you can compute one today with no tool.

The short version
AI answers have no positions to rank, so the only unit that survives re-running the question is a rate — how often you appear across many runs. You can compute one this afternoon with no tool. The catch is that a rate is a fraction, and whoever chose the questions chose your denominator, which is why two paid tools report two different numbers for the same brand and both are arithmetically correct.
Key highlights
- 1 A generated answer is not a list, so there is no position to report
- 2 The variation between runs is the signal, not noise on top of it
- 3 Your rate is appearances divided by runs, written as a fraction
- 4 Never average across engines — a blend hides the one that moved
- 5 Three runs and one hit is not 33%, it is one event
- 6 Ten runs of one question is the working floor for a rate
- 7 A vendor's prompt set is a vendor's denominator
- 8 Own and freeze your own question list before you buy anything
- 9 Logged out, one location, one time window, or it is a different experiment
- 10 Six lines beside the number make it reproducible by someone else
You ran your buying question through ChatGPT and you were named. You ran it again an hour later and you were gone. Nothing about your site changed in that hour.
The instinct is to decide which run was the real one. That instinct is the mistake, and it is the same one that makes two paid tools hand you two different numbers for the same brand in the same week.
There is no position to record, because a generated answer is not a list. What survives a re-run is a rate — how often you appear across many runs of the same question — and you can compute one this afternoon with no tool and no budget.
This piece is the method: what the unit is, how many runs it takes before the number means anything, who chose the denominator when a vendor hands you one, and what still contaminates the result after you have done everything right.
Why There is No Position to Record
A rank tracker answers one question: where does this URL sit in an ordered list of links. Every part of that sentence fails inside a chatbot answer.
The answer is generated, not retrieved from a fixed shelf. It names two or five or nine companies in a sentence, in an order that reflects how the sentence reads, not a scored placement. Run it again and the sentence is written again from scratch.
That is not a bug in the engine and it is not a bug in your site. It is what a language model does — the answer is composed at the moment you ask, out of whatever the model retrieved that time, in whatever order reads best in a paragraph.
Practitioners hit it constantly, and usually assume they have done something wrong:
10 people tested the same exact query at the same time. Only 3 got the same answer from ChatGPT.
So the variation is not noise sitting on top of the signal. The variation is the signal. Once you accept that, the measurement changes shape: you stop asking "where am I" and start asking "how often am I".
If you are still working out what counts as being named at all, what an AI mention actually is settles that first — a mention, a citation and a visit are three different objects.
The Unit That Survives a Re-run
Your rate is appearances divided by runs. One question, run N times, counting how many of those answers named you.
That is the whole unit. It is a percentage of runs, not a place in a list, and it holds up precisely because it was built out of the instability rather than in spite of it.
Three things it is not, and each one is a number somebody has tried to sell:
- Not a score out of 100. A score has already been through a formula you cannot see. A fraction has not.
- Not a share of voice. Share of voice compares you to a competitor set someone else assembled. Your rate compares you to your own last reading.
- Not one figure for "AI". It is one figure per engine, per question, per date. Collapsing those is where the number stops being true.
| What you were recording | What to record instead |
|---|---|
| position in the answer | did you appear at all — yes or no, per run |
| one screenshot | appearances ÷ runs, as a fraction you can show |
| "we rank #1 for X" | "we appeared in 6 of 10 runs of X, on 2026-09-09" |
| a single engine | the same fraction, kept separately per engine |
Two things follow from that table, and both are load-bearing.
A rate is only a rate with its denominator attached. "60%" means nothing. "6 of 10" can be checked, argued with, and re-run by the person you hand it to. Always write the fraction, never only the percentage.
Never average across engines. ChatGPT and Perplexity are not two samples of one population; they retrieve differently and reason differently. A blended score hides the one engine you are failing on, which is usually the one you would have acted on.
How Many Runs Before the Number Means Anything
Three runs and one hit is not 33%. It is one event with a percentage sign attached to it, and it will move ten points the next time you look.
The floor is uncomfortable and worth saying plainly. It is also where most of the market is quietly sitting:
AI share-of-voice extrapolates broad conclusions from a small number of prompts (usually 10-50). It's a good starting point, but it doesn't capture the full picture.
Two different sample sizes are in play here and they get confused constantly.
- Runs per question — how many times you ask the same question. This is what makes a single question's rate stable.
- Questions per claim — how many different questions you ask before saying anything about your category. This is what makes the claim generalise.
Ten runs of one question tells you about that question. Ten questions asked once each tells you about none of them. You need both, and the second is where the effort actually goes.
Below the floor, report a range or report nothing:
| Runs of one question | What you may honestly say |
|---|---|
| 1-2 | "we appeared" or "we did not" — one observation, no rate |
| 3-5 | a direction, stated as a range, with the raw count shown |
| 10 | a rate for that question, on that date |
| 10 across 10-20 questions | a rate for your category, on that date |
None of this requires a subscription. It requires you to run the same question ten times and write down ten yes-or-no answers, which is an afternoon.
The reason to do it by hand at least once is that you learn what the answers look like. A dashboard hands you 40% and no sense of whether those four hits were confident recommendations or a footnote inside a list of eleven. You cannot recover that from a number afterwards.
Two Tools, Two Numbers, and Neither of Them Lied
This is the part nobody in the category will tell you, because it is the part they sell.
A rate is a fraction, and somebody chose the denominator. When you buy a visibility tool, the vendor chose it — they assembled a synthetic prompt set, ran it, and reported the share of those prompts that named you.
A second vendor assembled a different set and reported a different share. Both numbers are arithmetically correct. They are answers to two different questions, and neither question was yours.
Practitioners work this out on their own, repeatedly:
The commercial tools all claim to solve this — they run synthetic prompt sets and report back your "citation share of voice" or "brand mention rate." I get the logic. But if the prompts are chosen by the vendor, and real buyers are asking for millions of slightly different variations with their own conversation histories, how representative is any of that data, really?
The fix is not a better vendor. The fix is that the prompt set is the asset and the score is the by-product. Own the list, and any tool becomes a way of running it faster. Rent the list, and you are renting your own denominator from someone whose product looks more necessary the lower your number goes.
Build the list before you measure anything
- Write down the questions a buyer actually asks before choosing. Not your keywords — the sentences. "Best X for a 20-person team", "X vs Y for someone on a budget", "is X worth it".
- Pull ten to twenty of them from real inputs. Sales call notes, support tickets, the questions in your own inbox. Search volume tells you nothing here; these questions are asked to a chatbot, not typed into Google.
- Freeze the list and date it. A prompt set that changes between measurements makes every comparison meaningless.
- Version it like a document. When you add questions, add them as v2 and keep reporting v1 alongside until you have enough runs on v2 to switch.
Do that and the number becomes yours. It also becomes comparable across months, which is the only comparison that actually exists — there is no public benchmark to compare it against, and anyone offering you one is offering you their sample.
Your Rate Belongs to a Test Account, Not to Your Brand
You checked from your laptop and you were named. Your client checked and you were not. Both of you are reading real output.
Answers are shaped by account history, location, session context and the model version you happened to get. So "our brand appeared for question X" is never the full sentence. The full sentence is "our brand appeared for question X, from this account, in this location, in this window".
That sounds like a reason to give up. It is the opposite — it is a list of controls, and they are all free:
- Logged out, or a clean account kept only for this. A personal account that has been discussing your industry for a year is not a neutral instrument.
- One stated location. Answers move across countries. Pick one, write it down, and keep it fixed.
- A stated time window. Run the ten runs inside one session, not scattered across a fortnight.
- The engine and the mode, named. "ChatGPT" is not specific enough when one mode searches the web and another answers from memory.
- The date, always. A rate with no date is not a measurement, it is an anecdote.
Change one of those and you have a different experiment, not a worse one. What you must not do is change one silently and then compare the two numbers.
What to Write Next to the Number
A rate that nobody else can reproduce is not evidence, and the difference between the two is about six lines of text.
Record this beside every number, every time:
- the question, verbatim
- the engine and mode
- runs, and appearances — as a fraction
- location and account state
- the date and the time window
- who ran it
That is the whole record. It fits in a spreadsheet row, it survives being forwarded to a client, and it is the thing that makes your number defensible when someone with a paid dashboard tells you a different one.
If you also want to see what the engines did to your site while you were measuring, your server log and GA4 already hold part of it — read that as a floor, never as a total.
Conclusion
There is no position to record, because a generated answer is not a list. What survives a re-run is a rate — appearances over runs — and it only means something with its denominator, its run count and its date attached. Own the question list, keep one figure per engine, and write down the conditions. That number is small and yours, which is more than most of what is being quoted at you.
Measuring AI Visibility: Frequently Asked Questions
What is AI visibility rate?
How many times should I run the same prompt?
Why do two AI visibility tools give me different numbers?
Should I average my score across ChatGPT, Perplexity and Gemini?
Does my location change my AI visibility?
Can I measure this without paying for a tool?
Find the exact growth leak in your business — in 2 minutes.
Paste your URL. Our AI agent crawls your site, diagnoses what's broken, and ships a step-by-step fix plan. Free, no signup.
Run free audit
