LLMs favour high-rated physicians. Brands, take note.
A 40,000-response audit shows LLMs weight star ratings and fees over nuance. Your brand's structured reputation footprint is now a visibility asset.
Key takeaways
- Raising a star rating from 3.9 to 4.7 increases LLM recommendation probability by 31.4 percentage points.
- Fee transparency is the second-largest driver: a $100 fee increase cuts recommendation probability by 20 points.
- LLMs act as infomediaries, picking up structured, numerical, widely-replicated signals rather than qualitative narrative.
- B2B brands with weak directory and review footprints will lose LLM visibility to competitors with stronger structured signals.
- Multilaterals and policy institutions should treat citation counts and indexed credentials as the equivalent of star ratings.
Reputation is doing most of the work. Across 40,068 scored responses, a single-variable change, raising a physician's star rating from 3.9 to 4.7, increased the probability of an LLM recommending that doctor by 31.4 percentage points. The arXiv preprint "Whose doctor does the AI recommend?", a preregistered randomised audit of seven models including GPT-4o-mini, makes that finding stark. No demographic sleight of hand is required to explain the pattern. Reputation signal dominates.
The audit methodology is worth understanding precisely because it is rigorous in a field saturated with anecdote. Researchers held gender and ethnicity constant via name-based correspondence-audit methodology, varied physician attributes independently across 3,024 choice sets, and ran nine prompt paraphrases across three patient personas. The result is a causal estimate, not a correlation. When you raise a rating, the model recommends more often. When you raise a fee from $90 to $190, choice probability falls by 20.0 percentage points. These are not small effects; together they account for a swing of more than 50 percentage points in recommendation probability from two variables alone.
What LLMs are actually reading
The implication for brand visibility is structural, not incidental. LLMs are not conducting fresh research when a user asks for a recommendation. They are pattern-matching against signals that appear prominently and repeatedly in their training data and, in retrieval-augmented systems, in indexed web content. Star ratings appear in schema markup, in Google Business Profiles, in review aggregators. Fee information appears in directories. Both signals are machine-readable, consistently formatted, and heavily replicated across sources. That is exactly the kind of signal a language model can anchor on.
This is the citation-pattern insight that the physician study inadvertently provides for any sector where LLMs are becoming recommendation intermediaries. The models are not reading qualitative prose and forming subtle judgements. They are picking up on structured, numerical, widely-replicated signals and weighting them heavily. A financial services firm asking why a competitor appears in ChatGPT's answer to "which wealth manager should I consider" should look first at its review footprint and fee transparency, not at the elegance of its thought-leadership content.
For multilateral institutions and UN agencies, the equivalent question is credibility proxies. These organisations rarely carry star ratings, but they do carry citation counts, co-publication records, and indexed appearances in high-authority directories. A World Bank affiliate that produces a report which is subsequently cited in sixty academic papers and indexed in three major policy databases has built exactly the kind of structured, replicable reputation signal that audit studies suggest LLMs favour. One that publishes internally and does not seed distribution has not.
The demographic finding is not a clean exoneration