LLMs over-cite popular papers and avoid critical citations
LLMs favour consensus over critique, creating a compounding citation disadvantage for organisations whose authority rests on original or contrarian research.
Key takeaways
- LLMs generate significantly fewer critical citations than human researchers across 132,000-plus citation contexts.
- Models over-cite popular and older papers; stronger models show this bias more, not less.
- LLMs cite authors socially closer to their training provenance, compounding visibility advantages for already-prominent institutions.
- Organisations producing contrarian or frontier research face a structural disadvantage in AI-mediated citation and retrieval.
- Citation mass and institutional recognition now matter more than argumentative quality for LLM visibility.
Across 132,000-plus citations drawn from 1,746 top NLP conference papers, a research team found that large language models systematically drain the critical edge from scientific literature. The arXiv preprint "Citing Less Critically" introduces a masked-citation task in which six popular LLMs each replace a human-written citation sentence with their own version, producing a counterfactual corpus that can be compared directly against what scholars actually wrote. The pattern that emerges is not random noise. It is structural.
Human researchers cite to argue. They use prior work as a foil, a counterexample, a limit case. LLMs cite to illustrate. Across all six models tested, critical citations, those that contrast or challenge a prior finding, appear significantly less often in LLM-generated text than in the human baseline drawn from the same 63,000-plus citation contexts. The models do not simply choose safer language around the same papers; they reach for different papers entirely.
What the models actually reach for
The second pattern reinforces the first. LLMs over-cite popular and older papers at a rate that exceeds human behaviour, and that tendency intensifies when the model is stronger. The implication is that capability and conformity travel together: a more powerful model is, on this measure, a more conservative one. It gravitates toward the canonical, the highly-cited, the already-settled. Papers at the frontier of a field, recent work still accumulating citations, work that challenges rather than consolidates, get systematically passed over.
The third pattern introduces a social dimension. Using a 20-million-edge coauthorship network, the researchers measured the social distance between an LLM and the authors it cites. Models cite closer to their own training provenance, the authors, institutions, and research clusters that appear more densely in their training data. That is a form of proximity bias with compounding effects: already-prominent researchers get cited more often, which improves their visibility in future training sets, which makes them more likely to be cited again.
Why this matters for organisations producing knowledge
For B2B organisations whose authority rests on intellectual output, the finding has a direct commercial read. Multilateral institutions, UN agencies, and policy bodies such as UNDRR or CGAP produce primary research precisely to shift how practitioners and policymakers reason about a problem. If LLMs flatten citation behaviour toward consensus and popularity, that research faces a structural disadvantage at the point of AI-mediated retrieval. A report that challenges an established framework, which is exactly the kind of work these institutions exist to produce, is less likely to appear in an LLM-generated literature review than a decade-old canonical paper from a more prominent institution.
Financial services firms and industrial groups that commission proprietary research face the same problem from a different angle. Thought leadership in these sectors tends to argue a position: a view on risk, a claim about market structure, a challenge to received methodology. LLMs trained to cite less critically will tend to strip that argumentative edge when they encounter the work, or skip it in favour of something blander and better-known.