AI visibility trackers disagree because none of them are “accurate”
AI visibility scores vary by methodology, not by error, so the only defensible number is the trend within one tool over time.
Key takeaways
- AI visibility platforms disagree because there is no fixed ground truth to measure against, unlike a calibrated instrument.
- Differences stem from methodology (prompt sets, sampling, citation definitions), not accuracy errors.
- Boards and auditors should be shown trend lines within one consistent tool, not absolute scores compared across vendors.
- Financial services firms and multilaterals citing AI visibility scores in public reports risk numbers that can't survive scrutiny.
- Running more queries does not fix the variance because LLM outputs are structurally unstable, not a sampling problem.
Two platforms, one client, one prompt set. The first tool reports the brand appears in 41% of relevant prompts. The second reports 23%. Somebody has to write a number in a board deck by Friday. Per iPullRank, this is now a weekly conversation, not an edge case.
The instinct is to ask which tool is wrong. That is the wrong question. iPullRank's argument, and it is the correct one, is that "accuracy" is not a concept that applies here at all. Accuracy implies a fixed, knowable ground truth against which a measurement can be checked, the way a thermometer's reading can be checked against a calibrated reference. There is no such reference for "does ChatGPT mention Holcim when asked about low-carbon cement." The answer depends on the exact prompt wording, the account's history, the model checkpoint, the day, sometimes the hour. Ask the same question twice in the same session and a non-trivial share of the time you get a different answer. The ground truth moves under your feet while you are trying to measure it.
What actually differs between platforms, then, is not accuracy but method: which prompts they sample, how many times they run each one, which models they query, whether they hold session state constant, how they define a "citation" (a link, a brand mention, a paraphrase of brand content). Two tools built on reasonable but different methodological choices will produce numbers that are both defensible and mutually irreconcilable. That is precision, not accuracy: each platform can be internally consistent and repeatable, and still disagree with a rival by a factor of two.
The number that matters is the trend, not the level
This has a specific consequence for anyone reporting AI visibility figures upward: stop presenting the level as if it were a fact and start presenting the direction as if it were the finding. A CMO who tells the board "we appear in 41% of prompts" has stated something that will be contradicted the moment procurement runs a competing tool. A CMO who says "our citation rate in this tool has risen from 18% to 31% over two quarters, measured consistently" has said something that survives scrutiny, because the methodology hasn't shifted, only the outcome has.
This is especially unforgiving territory for financial services firms and multilaterals, where numbers get audited, quoted in public filings, or cited in donor reports. A development bank citing its "AI visibility score" in an annual report is one methodology change away from a number that cannot be defended in front of a board or an auditor. The lesson from adjacent measurement disciplines, web analytics, media monitoring, applies directly: never let a single vendor's absolute number become the organisation's official metric. Treat it as one instrument's reading, useful for trend lines within itself, useless for comparison across instruments.
There is a second, quieter implication buried in iPullRank's framing: if the "ground truth" itself is unstable, sampling more is not a fix, it is a different kind of noise. Running a thousand queries against a model that gives probabilistically different answers to the same prompt does not converge on a stable brand-visibility percentage the way polling converges on an election result, because there is no fixed electorate. The variance is structural, not a sample-size problem. Vendors who imply otherwise, by promising a single, board-ready "AI visibility score," are selling false precision dressed up as accuracy.