voices quoted
Quoted in this piece
Why AI Visibility Tools Monitor but None of Them Fix
AI brand monitoring tools report where you stand but cannot tell you what to fix. What monitoring is genuinely worth paying for, and how to use it better.

The short version
AI visibility tools monitor well and cannot tell you what to fix, because monitoring is counting and fixing needs a causal claim nobody can yet evidence. Some of what a competitor wins is not a page you could have written at all. Monitoring is still worth paying for when it gives you a baseline, catches a drop and names who is recommended instead of you. Price the subscription against the hours it saves, not against insight it cannot deliver.
Key highlights
- 1 A tool describes a state; acting on it needs a change it cannot name
- 2 Monitoring needs no theory, fixing needs evidence that does not exist yet
- 3 A trustworthy fix needs a baseline, one change, a re-measure and a control
- 4 Some wins come from third-party mentions you cannot commission
- 5 Monitoring earns its cost on baselines, early drops and the competitor set
- 6 Pay for engine coverage and run volume, not the recommendation layer
- 7 Replace the vendor's prompts with yours and read the answers, not the score
- 8 Track the competitor column as your primary metric
- 9 The same limit applies to our own tooling
You bought the dashboard. It updates weekly, the charts are clean, and three months in nothing about your position has changed.
That is not a bad purchase. It is the category working exactly as built, and the gap between what it does and what you assumed it did is worth naming precisely — because once you see it, the tool becomes useful again for the thing it can actually do.
After the advent of LLM Search, I see dozens of tools offering LLM Monitoring. Sort of like "ahrefs, semrush for GEO". But when using these tools, I just surface-level information like "is your brand mentioned", "rank with respect to competitors", "most used sources" etc. I might be using these tools wrong, but are you ppl able to get any action-worthy data points from these tools?
Mentioned or not. Rank against competitors. Most-used sources. That is an accurate list of what the category delivers, and every item on it describes a state. None of them names a page to change.
Describing a State is Not the Same as Naming a Change
Hold the distinction and most of the frustration resolves.
| What the tool tells you | What you would need to act |
|---|---|
| you appeared in 34% of answers | which answers, and what the ones you lost had instead |
| a competitor outranks you | why that page was picked over yours, for that question |
| these are the most-cited sources | whether being on that source is achievable or even relevant |
| your score fell this month | whether anything you control caused it |
The right-hand column is not a harder version of the left. It is a different kind of claim. The left column reports what was observed. The right column asserts a cause — and that is the thing nobody in this market can currently demonstrate.
- Monitoring needs no theory. Run the query, record the answer, count. It is arithmetic and it is reliable.
- Fixing needs a causal claim. "Change this and the citation follows" requires evidence that the change moves the outcome, held against everything else that moved at the same time.
- That evidence does not exist yet. Engines update, indexes refresh, and competitors publish, all inside your measurement window. Isolating your own change is genuinely hard, not merely unfunded.
So the honest reading is not that these companies are lazy. They built the half that can be built without lying.
Why the Fix Half is the Harder Half
Consider what a trustworthy recommendation would require.
- A baseline — your rate on a fixed prompt set, stable enough that normal variation is known rather than assumed.
- One change, isolated, with nothing else shipped in the same window.
- A re-measurement long enough afterwards for the engines to have re-crawled and re-indexed.
- A control — questions you did not touch, to catch a category-wide move that would otherwise read as your success.
That is a controlled experiment, per recommendation, per customer. It is why the fix side is mostly generic best practice with a dashboard bolted on top: write clearly, answer early, add evidence. All good advice. None of it derived from your data, and none of it needing your subscription.
A vendor that hands you a confident, specific fix list is selling the half nobody has evidence for. That is worth being sceptical about even when the advice itself is fine — because the confidence is doing work the data cannot support.
Sometimes There is Nothing to Fix, Because It is Not Yours
The sharpest version of this problem is not that the tool won't tell you what to change. It is that the thing that worked was never a page you could have written.
One of our most-cited sources wasn't a page we wrote at all. It was a Reddit thread — 3 sentences, posted by someone I'd never heard of, mentioning our product in passing. That single thread got cited 19 times.
Eight months of logging, roughly fourteen hundred citations recorded, and the best-performing single source was three sentences by a stranger. One team, one account, unreplicated — but the shape of it is what matters.
- No content plan produces that. You cannot brief it, commission it or schedule it.
- No dashboard would have recommended it. The recommendation engine's whole input is your own properties.
- It is not an outlier you can dismiss either. Third-party mentions are a large share of what engines reach for, and most of them are outside your control by construction.
- It reframes what a low score means. Some of the gap between you and a competitor is not a content deficit. It is that somebody wrote about them somewhere and nobody wrote about you.
There is a version of acting on this, and it is not a content brief. If a stranger's post can outperform your own pages, the lever is being the kind of company people mention in those places — answering in communities where your buyers ask, being genuinely useful in a thread, having a product somebody wants to name. That work is slow, unschedulable and impossible to attribute, which is precisely why no dashboard recommends it.
- You can influence it and you cannot commission it. Showing up where the questions are asked raises the odds; it does not produce a deliverable.
- It compounds and it does not spike. Nothing about it will show up in next month's reading, which makes it hard to fund and easy to cut.
- It is the same work that has always produced word of mouth, now with a measurable downstream surface attached to it.
That last point is uncomfortable and it is the truthful one. A meaningful part of this metric is downstream of whether people talk about you, which is a much older problem than AI search and has no dashboard.
What Monitoring is Genuinely Good For
Having said all that: keep the tool, if the number is real. Monitoring earns its cost doing four things, none of which is telling you what to write.
- Establishing a baseline you can defend. A dated rate on a frozen prompt set is the thing every later conversation refers back to.
- Catching a drop early. You will not spot a fall from six-in-ten to two-in-ten by asking ChatGPT occasionally. A weekly run will.
- Naming the competitive set. Who gets recommended instead of you, per question, is the single most actionable output in the whole category — and it is monitoring, not fixing.
- Ending internal arguments. "Are we visible in AI search" becomes a number with a method behind it instead of an opinion contest.
None of those four requires the vendor to have an opinion about your content. They are all arithmetic on observations, which is the part this category does well.
Notice that the most useful item on that list is a by-product. The competitor names are not what the product markets itself on, and they are what you will actually act on.
What You are Actually Buying, Priced Honestly
Set against what it costs, the value is real but narrow. Worth seeing laid out.
| What you pay for | What you get | Could you do it by hand |
|---|---|---|
| engine coverage | the same questions run across four engines | yes, and it is the dullest hour of your month |
| run volume | ten runs per question instead of one | yes, and this is where the hours actually go |
| storage and history | twelve months of readings you did not have to keep | yes, in a spreadsheet, if you are disciplined |
| the recommendations | generic best practice, framed as personalised | you already have this for free |
| the interface | a chart somebody else will look at | this genuinely matters for client work |
Three of those five are hours you are buying back. One of them is presentation, which is a legitimate purchase if a client is reading it. One of them is the part to discount to zero.
- Price the subscription against the hours, not against the insight. If it saves six hours a month and costs less than six hours of your time, it pays for itself on that alone.
- Do not pay a premium for the recommendation layer. It is the cheapest thing in the box to produce and the most expensive to believe.
- Do pay for engine coverage if your buyers are spread across engines. That is a real cost base and the one thing genuinely hard to replicate manually.
Using the Dashboard You Already Pay For
Four changes make an existing subscription considerably more useful, and none of them requires a different vendor.
- Replace their prompt set with yours. If the tool allows it, upload real buying questions from sales calls. If it does not allow it, that is a finding about the tool.
- Read the answers, not the score. Once a month, open ten actual answers where you were absent and read what was said instead. That is where the next piece of work comes from.
- Track the competitor column as the primary metric. Your score moving tells you little; a named competitor appearing in four questions where you do not is a brief.
- Log what you shipped, with dates, in the same place. No one can hand you causation, but you can at least line your own changes up against the readings afterwards.
The fourth one is the closest anyone gets to closing the loop honestly, and it is a spreadsheet column rather than a feature. It will not prove causation. It will stop you from having no idea.
Which Applies to Us Too
It would be convenient to end by saying our own tooling solves this. It does not, and the same standard applies.
We can show you where you stand across the engines. We cannot tell you that changing a specific page will produce a specific citation, because that link has not been demonstrated by anyone, ourselves included. Anyone in this market claiming otherwise is asking you to accept a causal claim on the strength of a confident interface.
What is genuinely available today is smaller and real: a rate you can defend, a definition of what counts as a mention that you chose, the names of who is being recommended instead of you, and an honest read of what your own logs can and cannot see. That is a diagnostic. Treated as one, it is worth having.
Conclusion
Keep the dashboard if the number behind it is real, and stop expecting it to tell you what to write. Use it for a defensible baseline, early warning and a named competitor set, feed it your own questions, and log what you ship beside the readings so you can judge your own changes afterwards.
AI Visibility Monitoring Tools: Frequently Asked Questions
Do AI visibility monitoring tools actually help?
Why can't AI visibility tools tell me what to fix?
Should I cancel my AI visibility subscription?
What is the most useful output from these tools?
Can I get cited by a page I did not write?
Is monitoring worth paying for if it cannot fix anything?
Find the exact growth leak in your business — in 2 minutes.
Paste your URL. Our AI agent crawls your site, diagnoses what's broken, and ships a step-by-step fix plan. Free, no signup.
Run free audit

