BRP metric quantifies brand omission risk in LLM answers
Brand popularity no longer predicts LLM recommendation prominence. A new measurement framework makes the omission risk concrete.
Key takeaways
- Established brands are frequently omitted from LLM product recommendations, even when they are obvious category leaders.
- Recommendation prominence in LLMs correlates with broad marketplace presence across diverse sources, not brand popularity.
- Single-shot prompt tests measure noise; reliable LLM visibility requires repeated sampling across prompt variants.
- The BRP@k metric reframes brand absence in model outputs as measurable omission risk, not an unanswered query.
- The framework applies equally to commercial brands and institutions whose names should appear in policy or guidance answers.
Brand equity built over decades can vanish inside an LLM answer in a matter of tokens. A new paper published on arXiv by researchers studying brand and entity visibility in LLMs makes that risk measurable for the first time, introducing a framework that does not wait for a model to define who counts as a competitor.
The research names the core metric Brand Recommendation Probability, or BRP@k, where k is the rank depth examined. Pair it with Mean Reciprocal Rank and you have two numbers that describe not just whether a model mentions your brand, but where in the answer it appears. Across six LLMs and five product categories, the study finds that established brands are frequently omitted entirely, and that when brands do appear, prominence does not follow conventional brand-popularity rankings. Something else is driving recommendation order, and the paper points toward broader marketplace presence rather than share-of-market as the explanatory variable.
What the competitive-set problem reveals
Conventional search has always had a denominator: the index. A brand either ranks or it does not, and the universe of possible results is bounded. LLMs discard that structure. When a user asks for product recommendations in open-ended language, the model generates a candidate set from nothing, consulting no external list of eligible brands. The arXiv paper addresses this directly by defining the competitive set independently of model outputs. That is the methodological move that makes the framework genuinely useful: you decide which brands should be in contention, then measure how often the model agrees.
The consequence for measurement is significant. A brand that never appears in model outputs looks, under naive analysis, like it simply was not queried. The BRP framework reframes that absence as omission, a measurable gap between deserved presence and actual retrieval frequency. For a senior marketer at an industrial group or a financial services firm, the distinction is not semantic. Omission in an LLM answer is not the same as a bad organic ranking. A bad ranking can be worked on through conventional SEO. Omission means the model has, in effect, decided the brand is not part of the conversation.
Why popularity no longer predicts prominence
The finding that recommendation prominence does not follow brand popularity is the result that should trouble CMOs most. The intuitive assumption, one that has justified brand-building investment for generations, is that awareness creates preference creates recommendation. LLMs appear to break the middle link. A well-known brand in a category is not reliably surfaced ahead of a less prominent one.
What the paper associates with prominence instead is marketplace presence in a broader sense: the density and diversity of references across the text that made it into training data and retrieval corpora. This is not the same as brand awareness in any survey-research sense. A brand that features heavily in review aggregators, specialist forums, technical documentation, and editorial coverage across multiple contexts will be treated differently from one whose presence is concentrated in its own owned media or in a narrow channel.
For multilateral institutions and UN agencies, the implication lands differently but no less sharply. These organisations rarely think of themselves as competing for recommendation in a commercial sense, yet they are increasingly cited, or not cited, when LLMs answer policy questions, produce briefing summaries, or surface guidance on topics from climate risk to financial inclusion. The BRP framework is equally applicable to an institution whose name should appear in an answer about disaster risk reduction metrics as to a consumer brand in a product category.
Repeated sampling as the fix for stochastic outputs
One technical wrinkle the framework addresses is LLM non-determinism. Ask the same product question twice and the model may return different brands in a different order. Single-shot evaluation, the approach many brand-measurement vendors still use, produces a snapshot that could easily be an outlier. The arXiv researchers handle this through repeated sampling, estimating retrieval probability across multiple queries rather than recording a single output. The result is a distribution, not a point estimate, which is the correct way to characterise a stochastic system.
That methodological rigour has a direct commercial implication. Brands and their agencies that rely on one-off prompt tests to assess LLM visibility are measuring noise. A brand that appears in a manually run query may have a BRP of 0.12, meaning it surfaces in roughly one in eight equivalent prompts. The query that happened to return it tells you nothing useful about expected visibility.
Philanthropic and policy institutions that commission occasional AI audits face the same problem. A single favourable citation in a model output is not evidence of reliable presence. Reliable presence requires repeated sampling across prompt variants, which is what BRP@k operationalises.
The measurement gap that still needs closing
The framework is a research contribution, not yet a deployed product, and the paper's scope is five product categories across six models. Extrapolating to the full range of B2B categories, financial product queries, or institutional-knowledge domains requires further work. The association between marketplace presence and prominence is described but not fully explained: the mechanism by which broader textual presence translates into model retrieval probability is still a black box, which means the intervention pathway for a brand that scores poorly on BRP is inferential rather than proven.
That uncertainty, though, does not diminish the value of having a denominator at all. Brands that cannot quantify their LLM omission risk cannot prioritise the editorial, partnership, and content decisions that affect it. BRP gives them a number. The next question is which specific investments move it, and that is the measurement problem the industry will spend the next two years arguing about.