Decisions · Measurement

Change the Question, Change the Ranking

Ask an assistant the same buying question two normal ways and the brands it recommends barely overlap. The number on your AI visibility dashboard is mostly a property of the tracker, not of your business.

I went into this research expecting to find that AI visibility tools count brand mentions and nothing else. That the whole category sold a tally, the tally was thin, and the thinness was the story. I wrote that down before I searched, because a prediction you record in advance is the only kind that can embarrass you.

It came back false on the first search, and it was the prediction I held most confidently. The serious tools in this category separate a citation from a recommendation, on the explicit reasoning that a brand can be quoted as a source without ever being put forward as a provider. They track sentiment. They track position. They flag when an assistant describes the brand incorrectly. The sophisticated version of this product exists, and it is not expensive.

So the problem is not that the instrument is crude. The problem is what the instrument is pointed at. And if you are paying for one of these tools, or deciding whether to, that is the part that decides whether the money is well spent.

The same question, asked twice, two different shortlists

In May 2026, a team led by Will Jack published a study of production retrieval-augmented commercial recommendation. They ran roughly 6,000 paraphrase trials against roughly 6,000 same-prompt rerun controls on OpenAI and Anthropic models, and measured how much the recommended brand sets overlapped.

Rerunning an identical prompt produced sets that agreed about half the time. Rephrasing the same buying intent in a second natural way produced sets that agreed far less. A cosmetic rewording, "best CRM" against "top CRM", landed at 0.288. Adding a constraint a real buyer would add, "best CRM for a SaaS startup", dropped it to 0.135.

The Number
14–29%

Overlap between the brands an AI recommends when the same buying intent is phrased two natural ways, against 50–61% when the identical prompt is simply run twice. Raising the model's reasoning effort does not close the gap.

Source: Jack et al., "Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation," arXiv:2605.27440, May 2026

The authors do not leave the implication to the reader. Counting brand mentions across a fixed prompt set, they write, produces a metric whose dominant source of variance is which paraphrase the tracker happened to issue, rather than the model's behavior toward the brand. The prompt string, not the buyer's underlying intent, is the main input to which brands surface. They name this a direct challenge to a popular practice, and they are right to.

Read that once more with your own dashboard in mind. The thing moving your score most is a choice your vendor made about wording.

What you are actually paying to watch

Every measurement system has a division of labor between what the instrument decides and what the world decides. A good one keeps the instrument's contribution small. Here it runs the other way.

Decision Map · Whose choices are in the number
Decided by the trackerDecided by your business
Which prompts are in the libraryWhether the third-party record supports recommending you
How each buying intent is phrased — and whether a constraint is attachedWhat review sites, analysts and trade press say
How many paraphrases are issued per intentWhether your claims survive a check
Which assistants are polled, and how oftenWhat you sell, to whom, at what price
Whether answers shown inside a search results page count as AINothing in the left-hand column

The left column moves the score. The right column moves the business. The report hands you a single percentage and lets you assume it came from the right.

Two assistants, one brand, different answers

OpenAI and Anthropic models, measured side by side

May–June 2026
Case

A companion study ran the same prompt corpus and the same extraction pipeline across an OpenAI pool and an Anthropic pool. Across four independent measurement batches, the overlap between the two providers' recommendation sets sat at 0.349 to 0.356, with overlapping confidence intervals. The result replicated cleanly, and it fell below the rerun-stability baseline of 0.50 to 0.61 measured within a single provider.

The mechanism differs too. Brands recommended without appearing anywhere in the retrieval evidence, a pure-priors route, ran at 43–52% on the Anthropic flagships against 8–29% on the OpenAI cells.

Switching which assistant your tracker polls changes your brand set more than rerunning the same prompt does — so an engine-mix change in the vendor's roadmap arrives on your dashboard looking like a market move.

Nobody agrees how big this channel is either

Go looking for how much traffic AI assistants actually send, and you find a settled-sounding consensus that falls apart on contact.

Contentsquare's 2026 benchmark, built on 99 billion sessions across roughly 6,500 sites, puts AI-referred traffic at 0.2% of total visits, growing 632% year over year. Its own summary of that pair is the honest one: 0.2% is small and 632% is not.

Conductor's benchmark, across 1,215 enterprise domains and 3.3 billion sessions, puts the same quantity at 1.08%, with information technology reaching 2.80%.

Both are sound. They are answers to different questions being reported as one, and the gap between them is wider than most of the quarterly movements companies act on. The same collision shows up in conversion. Adobe measured AI-referred visitors to US retail sites converting 54% better than everyone else in May 2026; a year earlier the same measurement ran roughly half as well. A relative premium and an absolute rate are not the same object, and putting them on one slide produces a strategy that cannot be right.

Where this argument is weak

The load-bearing source is a preprint. It covers a specific set of models in a commercial retrieval setting, and Jaccard overlap on sets is a demanding measure that reads harsher than a looser one would. The authors report clustered bootstrap confidence intervals and say plainly that per-prompt precision is wide; the strength comes from averaging across many prompts, not from any single one.

The traffic evidence cuts against me as well as for me. A meaningful share of what analytics calls direct traffic is assistant-driven traffic that lost its referrer when somebody pasted a link into a fresh tab, and answers rendered inside a search results page are not separately attributed at all. The real footprint is larger than the low figures show. Small is not the same as ignorable, and a channel that converts well while measuring badly is the hardest kind to govern.

And the claim I most wanted to prove, I cannot. There is no public evidence that mention share fails to predict revenue. There is none that it succeeds, either. Nobody has run it, including the firms selling the dashboards. That is an open question, not a verdict, and anyone answering it confidently in either direction is selling something.

One more belongs here. Almost every non-academic source on this topic is published by a company that sells the category. Both findings this piece rests on came from arXiv, which is the only reason I trust them more than the rest.

When this is your problem
  • Buyers in your category actually ask an assistant who to hire — tested, not assumed.
  • The shortlist is contested — several credible providers, no default winner.
  • Somebody is reporting an AI visibility number upward to a board, an owner or a client.
  • And not yours if buyers arrive through a relationship, a distributor or a specification document.
  • And not yours if you have never established that anyone asks an assistant about your category. That test costs an afternoon and comes before buying anything.
  • And not yours if no decision would change on the answer. The subscription is not the cost. Spending against somebody else's prompt library is the cost.
The Path

Measure your own noise floor before you act on a single movement in the dashboard.

Take your ten highest-value buying questions. Write each one six to eight ways a real buyer would type it, including the constrained versions that name a size, a sector or a budget, because those are the phrasings the study found most volatile. Run all of them against the assistants your buyers use. Record which brands come back and how much the sets overlap. That overlap figure is your floor. Any movement in your visibility score smaller than your own floor is measurement, not news, and you decline to spend against it.

If your vendor cannot tell you how many paraphrases it issues per intent, that is the first question to ask, and the answer is diagnostic on its own.

What it costs
API credits and a spreadsheet. A few hundred dollars and one focused afternoon, or a day of an analyst's time if you want it repeatable.
How you'll know
Run it again a quarter later. A stable floor means you have a usable instrument and can start trusting movements above it. A floor that moves as much as your score means the dashboard is reporting its own configuration, and you stop paying for it.

One conversation. The number somebody is about to defend in a board meeting.

Not an audit of your vendor. Working out which of your reported metrics would survive being asked how it was measured.

Get a free positioning review Or read the frameworks →
This article is analysis for general information, not investment, financial or legal advice, and not a claim of wrongdoing by any company or person. Naming a company here is analysis, not accusation. Figures are drawn from published sources: Jack et al., "Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation," arXiv:2605.27440, May 2026; the companion cross-provider study, arXiv:2606.26116, June 2026; Contentsquare's 2026 Digital Experience Benchmark Report, based on 99 billion sessions across approximately 6,500 sites; Conductor's 2026 AI referral benchmark, across 1,215 enterprise domains and 3.3 billion sessions; and Adobe Analytics reporting on AI-referred retail conversion, May 2026. The two arXiv papers are preprints and are described as such in the text. Where a statement is the author's reasoning rather than a research finding, it is identified as such. Current as of September 2026. © 2026 CULT+MATH LLC.