# Your AI Visibility Report Changes Every Week
URL: https://doableclaw.com/blog/ai-visibility-tracking-clients/
> Weekly screenshots that contradict each other are under-sampled, not broken. Report the spread, and call the metric a diagnostic rather than a KPI.
Published: 2026-09-24

> **TL;DR:** Weekly screenshots that contradict each other are not a broken process, they are an under-sampled one — and the two look identical from the outside. The fix is to stop defending a single reading and report the spread instead. Then describe the metric honestly: it is a diagnostic, not a KPI, because it has neither a benchmark nor a stable denominator.

Somebody on your team checks the prompts by hand every Monday, screenshots the answers, and pastes them into the client deck. By Thursday the answers are different. The client notices before you do.

> Someone on the team manually checks prompts in ChatGPT and Perplexity each week then adds screenshots to the client report. It takes too long and the data feels hard to defend because the results can change between checks.
>
> — r/DigitalMarketing · [Reddit](https://www.reddit.com/r/DigitalMarketing/comments/1v1tibl/our_ai_visibility_reports_feel_made_up_am_i_the/)

Twenty-two replies, most of them sympathetic. But it is not made up. It is under-sampled, which looks identical from the outside and has a completely different fix.

This piece is how to stop defending a screenshot, what to put in front of the client instead, and how to describe what this metric is without promising something you cannot deliver.

## The Quick Answer

- [One screenshot is a sample of one, and it will contradict itself](#not-the-screenshot)
- [A screenshot is illustration underneath the number, not the measurement](#not-the-screenshot)
- [Report 6 of 10 runs, with last month's fraction beside it](#the-spread)
- [Volatility is reportable, and clients accept it faster than agencies expect](#the-spread)
- [A KPI needs a benchmark and a stable denominator; this has neither](#not-a-kpi)
- [Never promise it becomes a KPI on a timeline](#not-a-kpi)
- ["Our other agency's tool says 40%" — different prompt set, different denominator](#pushback)
- [Claiming revenue attribution is the fastest way to lose the account](#pushback)
- [Twenty accounts by hand each week is 200 checks and the least defensible artefact](#the-cost)
- [Monthly with ten runs is usually less work than weekly single checks](#the-cost)

## The screenshot was never the artefact

A screenshot is one run. One run of a generated answer tells you what happened that time, and the next time is written from scratch.

That instability is not a defect in the tool, the prompt, or the person doing the checking. A language model composes the answer at the moment you ask, out of whatever it retrieved on that attempt, in whatever order reads best in a paragraph. Two people asking the same question in the same minute can get different answers.

So the weekly ritual has a structural problem, not a diligence problem:

* **One capture per week is a sample of one.** You are not tracking a trend, you are collecting anecdotes at a fixed interval.
* **It contradicts itself, publicly.** Last week's screenshot showed the client named. This week's does not. Both are true, and side by side they read as unreliability.
* **It cannot show a direction.** You cannot tell an improvement from noise without knowing how much the number moves on its own.
* **It costs the most and proves the least.** Manual weekly checks across twenty accounts is real payroll spent producing the least defensible artefact in the deck.
* **It invites the wrong question.** A screenshot makes the client ask "why did this change", which you cannot answer. A rate makes them ask "is it going up", which you can.

The screenshot still has a job. It is illustration — proof that the answer looked like this, sitting underneath the number. It is not the measurement.

There is one more thing worth carrying into the report while you are at it, and it costs nothing: referral traffic from the AI hosts. [It is sitting in GA4 and in your server log already](/blog/check-ai-traffic-free), and unlike the rate it is a hard count of sessions that genuinely arrived. Report it as a floor rather than a total — it can only ever see the mentions that carried a link and got clicked — and label it that way, because a client who mistakes it for the whole picture will conclude the work is not working.

## Report the spread, not the reading

The fix is to run the same question enough times that the variation becomes visible, and then to show the variation instead of hiding it.

| What you show today | What to show instead |
|---|---|
| one screenshot, weekly | the same question run ten times, monthly |
| "the client appeared" | "named in 6 of 10 runs" |
| a new capture each week that disagrees | last month's fraction beside this month's |
| an implied stability nobody has | the spread, stated as the finding |

Once the number is a fraction, the thing that used to embarrass you becomes the most interesting line in the report. "This question is volatile — we were named in 6 of 10 runs, and last month it was 4 of 10" is a sentence with information in it. "Here is Monday's screenshot" is not.

**Volatility is itself reportable, and clients understand it faster than agencies expect.** Anyone who has run paid media knows what a noisy metric looks like. What they will not forgive is being shown a single reading presented as a settled fact, and then being shown a different one a fortnight later.

Two practical moves make the spread readable:

1. **Group the runs, do not average them away.** Six named, three absent, one cited-without-being-named is a more useful line than "60%".
2. **Keep the run count constant.** If it was ten last month it is ten this month. Changing the sample size between reports breaks the comparison as surely as changing the questions.

site: doableclaw.com/ai-visibility — Put a rate and a competitor list in front of the client instead of a screenshot.

---

## Say out loud that it is not a KPI yet

The other half of defending the number is describing it honestly, and practitioners are already asking the right question:

> I'm building some internal tooling around this and wondering if I'm overthinking it or if this becomes a legit KPI in 12–24 months.
>
> — r/SEO · [Reddit](https://www.reddit.com/r/SEO/comments/1otf9z1/anyone_here_actively_tracking_llm_citations_as_a/)

Here is the honest answer. A KPI needs two things this metric does not have: a benchmark to judge the number against, and a stable denominator so the number means the same thing next quarter. There is no credible public baseline for a good mention rate, and the denominator is whichever prompt set someone chose.

That does not make it useless. It makes it a diagnostic.

| | A KPI | A diagnostic |
|---|---|---|
| what it answers | are we hitting the target | what is happening, and where |
| needs a benchmark | yes | no |
| judged against | an external standard | its own previous reading |
| what it drives | accountability | the next piece of work |
| safe on a retainer | when it is mature | when it is labelled honestly |

Report it as a diagnostic and it survives scrutiny, because you are not claiming more than you have. Report it as a grade out of 100 and the first person who asks what the denominator is will take the whole deck apart.

* **Do not promise it will become a KPI on a timeline.** Nobody knows, and a promised timeline is a hostage.
* **Do not attach it to revenue.** A mention influences a decision you cannot observe. Any line connecting this to pipeline is a guess wearing a chart.
* **Do label it, in the report, every month.** One sentence: this is a diagnostic against our own frozen question list, not an industry benchmark.

## What the client actually pushes back on

Four objections come up, and each has a short answer that does not require you to overclaim.

| What they say | What to answer |
|---|---|
| "this changed since last time" | one run is a sample of one — here is the rate across ten, and here is last month's |
| "our other agency's tool says 40%" | different prompt set, different denominator; both correct, neither comparable |
| "is this making us money" | not measurable yet, and we will not pretend otherwise — it tells us where we are absent |
| "why should I pay for a number that moves" | because it moves whether or not you measure it, and being absent is the finding |

The third one is the hardest and the most important to get right. **An agency that claims revenue attribution for AI mentions is writing a cheque it cannot cash**, and the day it gets audited is the day the retainer ends. Saying "we cannot attribute this yet, and here is what it does tell you" is a stronger position than it feels like at the time.

The fourth answer is the one that renews retainers. The client's absence from an answer is happening now, to real buyers, whether anybody is counting it or not. Measuring it does not cause it. It just means somebody noticed.

## What the weekly ritual is actually costing

Before defending the artefact, it is worth pricing it. Manual weekly checking across a book of accounts is one of the more expensive things an agency does without noticing.

* **It scales linearly with clients.** Twenty accounts, five prompts each, two engines, every week. That is two hundred manual checks a week, and the number only goes up as you grow.
* **It is the least delegable hour in the week.** Someone has to read each answer and decide whether the mention counted, and that judgement is exactly what nobody wrote down.
* **The output has the shortest shelf life in the deck.** A screenshot is stale the day after it is taken, and it will be contradicted before the report is even sent.
* **It produces no history you can use.** Fifty-two screenshots is not a time series. Twelve monthly rates is.
* **It trains the client to ask about the wrong thing.** Every capture invites a conversation about that one answer instead of about the trend.

Monthly, with ten runs per question, is usually *less* total work than weekly single checks — and it produces something that can go in front of a finance director. The reason most teams have not switched is that the weekly ritual feels more diligent. It is not; it is more frequent, which is a different thing.

**Decide what counts as a hit before anybody runs anything.** Named in the conclusion with a reason, listed among options, and cited without appearing in the text are three separate outcomes. Write the rule at the top of the sheet, apply it every month, and the judgement stops being an argument.

## What this looks like on the renewal call

The version that survives is unglamorous and specific:

* **One frozen list of the client's real buying questions**, which you own and they can read.
* **A fraction per question per engine**, with the run count and the date beside it.
* **The names of who was recommended instead**, which is the part that stops a meeting.
* **Last month's figures next to this month's**, because that is the only benchmark that exists.
* **One honest sentence about what it is** — a diagnostic, not a KPI, not attributed to revenue.
* **A screenshot or two underneath**, as illustration, clearly labelled as one run.

None of that requires a tool. It requires the discipline to run the same questions the same way every month and to resist summarising them into a single number, which is the one thing every dashboard is built to do for you.

One refinement worth adding once the basics hold. Being named and being recommended are not the same outcome, and [the gap between them is where most of the disappointment lives](/blog/named-is-not-recommended) — a client who appears in every lineup and is never the recommendation has a different problem from one who is absent. Counting those separately turns a flat rate into a piece of direction.

## Bottom Line

The results changing between checks is the finding, not the failure. Show a rate across a fixed number of runs, with last month beside it and the competitor names attached, and label it for what it is — a diagnostic against your own frozen question list, not an industry benchmark and not attributable to revenue. That version survives the renewal call. A confident grade out of 100 does not.

## Defending an AI visibility report: frequently asked questions

### Why do my AI visibility results change between checks?

Because a generated answer is composed fresh each time from whatever the model retrieved on that attempt. Two identical questions minutes apart can return different companies. One check is a sample of one, so the change is expected rather than a sign that something broke.

### How do I explain changing AI results to a client?

Stop defending a single reading. Show the rate across a fixed number of runs, with last month's rate beside it, and say plainly that the variation is part of the measurement. Clients accept a noisy metric that is described honestly far more readily than a stable-looking one that contradicts itself.

### Is AI visibility a real KPI?

Not yet. A KPI needs a benchmark and a stable denominator, and this metric has neither — there is no credible public baseline, and the denominator is whichever prompt set was chosen. Report it as a diagnostic judged against the client's own previous reading.

### Should I put AI mentions in the same report as SEO rankings?

Yes, as its own row, never merged into the ranking rows. They measure different things — one counts clicks that arrived, the other counts being named in an answer that may produce no click at all.

### Can I attribute revenue to AI mentions?

No, and claiming you can is the fastest way to lose the account when someone checks. A mention influences a decision that leaves no trace in your analytics. Report it as a leading diagnostic and keep revenue attribution to the channels that can actually carry it.

### How often should I run the checks?

Monthly, with enough runs to produce a rate, beats weekly single screenshots on both cost and defensibility. If something specific changes mid-month — a site migration, a big content push — run an extra set and label it separately rather than folding it into the monthly series.
