AE

Quoted in this piece

What Counts as an AI Mention? Three Metrics, Not One

One mention rate is three events added together: recommended, listed, cited-only. Count them separately, and treat a zero as a broken instrument first.

A mention rate of 7 of 10 splitting into recommended 2, listed 4 and cited only 1

The short version

A single "mention rate" is three different events added together: being recommended with a reason, being listed among options, and being cited with your name nowhere in the answer. A buyer does not experience those the same way, so a blended figure cannot move for any reason you can act on. And a zero is more often a broken instrument than a real finding.

Your tool reports a 40% mention rate. Forty percent of what, exactly?

Somewhere underneath that figure, three different events were counted as the same thing: an answer that told the buyer to choose you, an answer that listed you among four options, and an answer that never said your name but attached your link at the bottom. Those are not one metric. A buyer does not experience them as one metric either.

Measuring AI visibility has two halves. One is how to count — you run the same question many times and record how often you appear, because a generated answer has no positions to rank. The other is this one: what you are allowed to count as a hit in the first place, and how to tell a real zero from an instrument that quietly stopped working.

If the first thing you need is the definition — a mention, a citation and a visit are three different objects — start there and come back.

The Number is Three Numbers Stacked

Practitioners worked this out before the vendors did:

"Mention rate" as a single number is kind of a trap. It's not one thing — it's at least three different things we're lumping together and calling one metric.
r/aeo · Reddit

The three, in descending order of what they are actually worth to you:

TierWhat the answer didWhat it is worth
Recommendednamed you in the conclusion and said whythe buyer has a reason to pick you
Listedput you in a lineup of three or four, no ranking, no reasonyou are in the consideration set
Cited onlynever said your name, attached your linkyou were read, not recommended

Read the right-hand column as three separate outcomes, because that is what they are. A recommendation ends a decision. A listing enters you into one. A bare citation means the engine used your page to answer a question about somebody else.

We have written the qualitative version of this before — being named is not the same as being recommended walks through why the gap exists and where in a conversation you tend to fall out of it. This piece is the measurement half of that argument.

Why Adding Them Together Breaks the Report

Adding the three tiers and dividing by runs produces a number that cannot move for any reason you can act on.

Say you run twenty tests and land ten hits. If those ten are all bare citations, you have a content problem: the engine reads you and does not think you are the answer. If they are all recommendations, you are winning and should be spending on the questions you are not testing yet. Both report as 50%.

The blended number also hides the direction you are moving in. A quarter where five recommendations decay into five listings reads as flat. It is not flat. It is the early half of losing the category, and the only instrument that would have shown it is the one you collapsed.

There is a third failure that is worse, because it looks like success:

  • Citations inflate easily. A page that answers a comparison question well gets attached to answers about your competitors. Your citation count rises while your recommendation count does not move.
  • Listings inflate on volume. Ask broader questions and you appear in more lineups. Ask the questions a buyer actually asks and the lineups get shorter.
  • Recommendations are the only tier that resists being gamed by asking easier questions, which is exactly why a vendor's blended figure is usually carried by the other two.

So the rule is short. Three tiers, three counts, three fractions, reported side by side. If you must have one headline number, make it the recommendation rate and show the other two beneath it.

What that changes in practice:

  • Your target stops being a percentage and starts being a tier. "Move six listings into recommendations" is work somebody can do. "Get to 55%" is not.
  • A flat quarter becomes readable. Same headline, three different stories underneath, and only one of them means do nothing.
  • You can tell a content problem from a positioning problem. High citations and low recommendations means the engine reads you and does not rate you. That is a different fix from not being retrieved at all.
  • The report survives a sceptical reader. Anyone can ask "what counted as a mention" and you have an answer that fits on one line.

How to Count Each Tier Without Arguing About It

Most of the disagreement about what counts as a hit disappears once you write the test down before you run it.

  1. Recommended — your name appears in the answer's conclusion or its direct advice, with a stated reason. "Go with X if you need Y" counts. A closing sentence that lists four names does not.
  2. Listed — your name appears in the body of the answer as one option among others, with no reason attached to you specifically.
  3. Cited only — your URL appears in the sources and your name appears nowhere in the answer text.
  4. Absent — neither. Record what the answer said instead; the names that did appear are your real competitive set for that question.

Two judgement calls come up every time, so decide them once:

  • A name in the conclusion with no reason counts as Listed, not Recommended. The reason is what makes it a recommendation.
  • A name that appears with a caveat attached — "X is cheaper but harder to set up" — still counts as Recommended. It gave the buyer a reason, and a mixed reason is a reason.

Write your two calls at the top of the sheet. Anyone re-running your measurement has to make the same ones, and a rate that two people can produce independently is the only kind worth putting in front of a client.

Free tool
See which answers recommend you, which only list you, and which just cite the page.
Check AI visibility

Zero is a Reading, and Usually a Broken Instrument

You ran the set and got nothing. Before you take that to anyone, assume you broke it.

A true zero is real and it happens to strong brands. Being well known offline, being the biggest name in a category, and having a site that ranks well do not put you inside a generated answer — the engine is answering the question in front of it out of what it retrieved, not out of what everybody knows. A brand can be famous and absent at the same time, and finding that out is worth the afternoon.

But a broken run produces exactly the same figure, on the same dashboard, in the same colour. A zero is a broken instrument until you have proved otherwise. That is our standing rule for our own work, and it has caught more of our own mistakes than any other check we run.

There is a practical reason to be strict about this beyond being right. A zero is the one reading you will not be allowed to soften. Take a real one to a client and the conversation is about what to do; take a broken one and you have spent your credibility on a bug, and the next real finding you bring gets discounted.

The things that produce a false zero are dull and common:

  • The prompt never ran. A timeout, a rate limit, or a session that expired halfway through the set.
  • The engine answered from memory. Some modes search the web and some do not. A mode that did not search cannot cite anyone.
  • The question was too narrow to have an answer. If nobody is named, you have measured the question, not yourself.
  • Your brand name is ambiguous. The engine named a different company with your name, and your string match missed it.
  • The matching is too strict. "Acme" and "Acme Inc." and "acme.com" are the same hit, and a naive check counts two of them as misses.

The Test That Separates the Two

This costs one minute and it settles it every time.

  1. Look at whether anybody was named. A true zero still shows your competitors' names in the answer. A broken run shows nothing for anyone, or shows the same nothing on every question in the set.
  2. Re-run three of the questions by hand. Not through the tool. Open the engine, paste the question, read the answer yourself.
  3. Check one question you know you win. Keep a control question in every set — one you are confident about. If the control comes back zero, the instrument is broken, not the brand.
  4. Compare against the last run's competitor list. If the names that used to appear have also vanished, the problem is upstream of you.

If it survives all four, the zero is real, and it is the most useful reading in the report. It is also a much easier thing to explain to a client than a soft 12%, because there is nothing to argue about — the answer named four companies and you were not one of them.

What a Defensible Hit Record Looks Like

The output of all of this is not a score. It is a small table somebody else can re-run.

FieldExample
question, verbatim"best expense tool for a 20-person team"
runs10
recommended2
listed4
cited only1
absent3
who else was namedthe four names, listed
engine, mode, location, dateChatGPT, web search on, UK, 2026-09-09

Four things stay out of that record, and each one has cost somebody a bad quarter:

  • No blended percentage. If a single number is unavoidable, it is the recommendation rate, labelled as such.
  • No comparison to another brand's published figure. Different prompt set, different denominator, not comparable.
  • No sentiment score. You are counting whether you were named and why. Grading the tone of the answer is a different project and a much softer one.
  • No month-on-month delta from a set that changed. If the questions moved, the comparison is void — start the series again and say so.

Notice that the four counts add up to the run count. That is the check: if they do not, a run was lost and the rate underneath is wrong. Notice also that the competitor names are part of the record, not a footnote — they are the only part of this a buyer would recognise.

You can see part of what the engines did to your site while you were measuring in tools you already own, and your server log and GA4 hold that half. Read it as a floor, never as a total: a mention with no link attached leaves no trace there at all.

Conclusion

Stop reporting one mention rate. Count recommendations, listings and bare citations separately, report three fractions side by side, and if you must have a headline number make it the recommendation rate. Before you take a zero anywhere, prove the instrument was working — a true zero still names your competitors, and a broken run names no one at all.


What Counts as an AI Mention: Frequently Asked Questions

What is a good AI mention rate?
Nobody can honestly tell you. No credible public baseline exists, and a vendor's benchmark is a vendor's sample. Measure yourself against your own last reading, per tier, with the date attached. That comparison is real because you generated both halves of it.
Should a citation count as a mention?
Count it, but never in the same column as a recommendation. A citation means the engine read your page and used it to answer a question — often a question about somebody else. It is evidence you are retrievable. It is not evidence you are the answer.
Why does my mention rate differ from my agency's?
Almost always because one of you is counting a different set of tiers as hits, or running a different prompt set. Compare the two definitions before you compare the two numbers — in our experience the definitions disagree more often than the measurements do.
My AI visibility is zero. Is that possible?
Yes, and it is common for well-known brands. But check the instrument first: a true zero still shows your competitors' names in the answer, while a broken run shows nothing for anyone. Keep one control question you expect to win, and use it to tell the two apart.
How do I count a mention with my brand spelled differently?
As a hit. Decide your accepted variants — the name, the legal name, the domain, the common misspelling — write them at the top of the sheet, and apply them to every run. A matcher that is stricter than a human reader will manufacture zeros.
Do I need a tool to track this?
No, not to start. Ten runs of one question, four tallies and a date is a real measurement, and doing it by hand once teaches you what the answers actually look like. A tool buys speed after your prompt set and your tier definitions exist — it cannot supply either.

Find the exact growth leak in your business — in 2 minutes.

Paste your URL. Our AI agent crawls your site, diagnoses what's broken, and ships a step-by-step fix plan. Free, no signup.

Run free audit