Change the Question, Change the Ranking
Ask an assistant the same buying question two normal ways and the brands it recommends barely overlap. The number on your AI visibility dashboard is mostly a property of the tracker, not of your business.
I went into this research expecting to find that AI visibility tools count brand mentions and nothing else. That the whole category sold a tally, the tally was thin, and the thinness was the story. I wrote that down before I searched, because a prediction you record in advance is the only kind that can embarrass you.
It came back false on the first search, and it was the prediction I held most confidently. The serious tools in this category separate a citation from a recommendation, on the explicit reasoning that a brand can be quoted as a source without ever being put forward as a provider. They track sentiment. They track position. They flag when an assistant describes the brand incorrectly. The sophisticated version of this product exists, and it is not expensive.
So the problem is not that the instrument is crude. The problem is what the instrument is pointed at. And if you are paying for one of these tools, or deciding whether to, that is the part that decides whether the money is well spent.
The same question, asked twice, two different shortlists
In May 2026, a team led by Will Jack published a study of production retrieval-augmented commercial recommendation. They ran roughly 6,000 paraphrase trials against roughly 6,000 same-prompt rerun controls on OpenAI and Anthropic models, and measured how much the recommended brand sets overlapped.
Rerunning an identical prompt produced sets that agreed about half the time. Rephrasing the same buying intent in a second natural way produced sets that agreed far less. A cosmetic rewording, "best CRM" against "top CRM", landed at 0.288. Adding a constraint a real buyer would add, "best CRM for a SaaS startup", dropped it to 0.135.
Overlap between the brands an AI recommends when the same buying intent is phrased two natural ways, against 50–61% when the identical prompt is simply run twice. Raising the model's reasoning effort does not close the gap.
The authors do not leave the implication to the reader. Counting brand mentions across a fixed prompt set, they write, produces a metric whose dominant source of variance is which paraphrase the tracker happened to issue, rather than the model's behavior toward the brand. The prompt string, not the buyer's underlying intent, is the main input to which brands surface. They name this a direct challenge to a popular practice, and they are right to.
Read that once more with your own dashboard in mind. The thing moving your score most is a choice your vendor made about wording.
What you are actually paying to watch
Every measurement system has a division of labor between what the instrument decides and what the world decides. A good one keeps the instrument's contribution small. Here it runs the other way.
| Decided by the tracker | Decided by your business |
|---|---|
| Which prompts are in the library | Whether the third-party record supports recommending you |
| How each buying intent is phrased — and whether a constraint is attached | What review sites, analysts and trade press say |
| How many paraphrases are issued per intent | Whether your claims survive a check |
| Which assistants are polled, and how often | What you sell, to whom, at what price |
| Whether answers shown inside a search results page count as AI | Nothing in the left-hand column |
The left column moves the score. The right column moves the business. The report hands you a single percentage and lets you assume it came from the right.
Two assistants, one brand, different answers
OpenAI and Anthropic models, measured side by side
May–June 2026A companion study ran the same prompt corpus and the same extraction pipeline across an OpenAI pool and an Anthropic pool. Across four independent measurement batches, the overlap between the two providers' recommendation sets sat at 0.349 to 0.356, with overlapping confidence intervals. The result replicated cleanly, and it fell below the rerun-stability baseline of 0.50 to 0.61 measured within a single provider.
The mechanism differs too. Brands recommended without appearing anywhere in the retrieval evidence, a pure-priors route, ran at 43–52% on the Anthropic flagships against 8–29% on the OpenAI cells.
Switching which assistant your tracker polls changes your brand set more than rerunning the same prompt does — so an engine-mix change in the vendor's roadmap arrives on your dashboard looking like a market move.
Nobody agrees how big this channel is either
Go looking for how much traffic AI assistants actually send, and you find a settled-sounding consensus that falls apart on contact.
Contentsquare's 2026 benchmark, built on 99 billion sessions across roughly 6,500 sites, puts AI-referred traffic at 0.2% of total visits, growing 632% year over year. Its own summary of that pair is the honest one: 0.2% is small and 632% is not.
Conductor's benchmark, across 1,215 enterprise domains and 3.3 billion sessions, puts the same quantity at 1.08%, with information technology reaching 2.80%.
Both are sound. They are answers to different questions being reported as one, and the gap between them is wider than most of the quarterly movements companies act on. The same collision shows up in conversion. Adobe measured AI-referred visitors to US retail sites converting 54% better than everyone else in May 2026; a year earlier the same measurement ran roughly half as well. A relative premium and an absolute rate are not the same object, and putting them on one slide produces a strategy that cannot be right.
Where this argument is weak
The load-bearing source is a preprint. It covers a specific set of models in a commercial retrieval setting, and Jaccard overlap on sets is a demanding measure that reads harsher than a looser one would. The authors report clustered bootstrap confidence intervals and say plainly that per-prompt precision is wide; the strength comes from averaging across many prompts, not from any single one.
The traffic evidence cuts against me as well as for me. A meaningful share of what analytics calls direct traffic is assistant-driven traffic that lost its referrer when somebody pasted a link into a fresh tab, and answers rendered inside a search results page are not separately attributed at all. The real footprint is larger than the low figures show. Small is not the same as ignorable, and a channel that converts well while measuring badly is the hardest kind to govern.
And the claim I most wanted to prove, I cannot. There is no public evidence that mention share fails to predict revenue. There is none that it succeeds, either. Nobody has run it, including the firms selling the dashboards. That is an open question, not a verdict, and anyone answering it confidently in either direction is selling something.
One more belongs here. Almost every non-academic source on this topic is published by a company that sells the category. Both findings this piece rests on came from arXiv, which is the only reason I trust them more than the rest.
- Buyers in your category actually ask an assistant who to hire — tested, not assumed.
- The shortlist is contested — several credible providers, no default winner.
- Somebody is reporting an AI visibility number upward to a board, an owner or a client.
- And not yours if buyers arrive through a relationship, a distributor or a specification document.
- And not yours if you have never established that anyone asks an assistant about your category. That test costs an afternoon and comes before buying anything.
- And not yours if no decision would change on the answer. The subscription is not the cost. Spending against somebody else's prompt library is the cost.
Measure your own noise floor before you act on a single movement in the dashboard.
Take your ten highest-value buying questions. Write each one six to eight ways a real buyer would type it, including the constrained versions that name a size, a sector or a budget, because those are the phrasings the study found most volatile. Run all of them against the assistants your buyers use. Record which brands come back and how much the sets overlap. That overlap figure is your floor. Any movement in your visibility score smaller than your own floor is measurement, not news, and you decline to spend against it.
If your vendor cannot tell you how many paraphrases it issues per intent, that is the first question to ask, and the answer is diagnostic on its own.
- What it costs
- API credits and a spreadsheet. A few hundred dollars and one focused afternoon, or a day of an analyst's time if you want it repeatable.
- How you'll know
- Run it again a quarter later. A stable floor means you have a usable instrument and can start trusting movements above it. A floor that moves as much as your score means the dashboard is reporting its own configuration, and you stop paying for it.
One conversation. The number somebody is about to defend in a board meeting.
Not an audit of your vendor. Working out which of your reported metrics would survive being asked how it was measured.
Get a free positioning review Or read the frameworks →