Most LLMs attach quotes to 90% of claims. Accuracy is another matter.
High citation rates mask a substantiation gap that directly threatens the credibility of any institution whose guidance appears in AI-generated answers.
Key takeaways
- Most LLMs attach verbatim quotes to over 90% of factual claims in clinical QA from prompting alone.
- Fewer than half those quotes fully substantiate the claims they are meant to support.
- The substantiation failure persists across all 12 models tested and cannot be fixed by prompting.
- Institutions whose documents are cited by LLMs may receive citation credit for claims they would dispute.
- Being cited by an LLM and being cited accurately are not the same objective; brands need to audit both.
Twelve LLMs, 222 clinical questions, four sets of practice guidelines: the arXiv study "Verifiable by Construction" is the most granular public audit of citation behaviour in clinical question answering yet published, and its headline finding is not reassuring. Most models attach verbatim quotes to more than 90% of their factual claims. Fewer than half those quotes actually substantiate the claims they are meant to support.
That gap, between the appearance of citation and the reality of verification, is the central problem for any organisation whose credibility depends on what AI systems say about it.
Citation as theatre
The mechanics are worth stating plainly. In clinical QA, a model is asked a factual question, retrieves passages from reference material such as practice guidelines, and attaches a quoted excerpt to each claim it makes. The quote is verbatim: lifted directly from the source, not paraphrased. The visual effect is of a properly footnoted answer. The epistemic effect is often something else entirely.
The arXiv paper finds that the failure occurs in the third step: the quote exists, it is real text from the source document, but it does not fully substantiate the specific claim it is supposed to anchor. A model might claim that a drug is contraindicated in renal failure and attach a quote discussing renal dosing adjustments. The quote is related; it is not sufficient. For a clinician working under time pressure, that distinction is invisible unless they open the original document, which is precisely what verbatim citation was supposed to make unnecessary.
The study measures three stages separately: citation coverage (does every claim get a quote?), verbatim accuracy (is the quote faithful to the source?), and substantiation (does the quote actually prove the claim?). Coverage approaches 90% across most of the twelve models tested. Verbatim accuracy is generally high. Substantiation is where the system collapses.
What this means for brands that are cited, or want to be
The clinical setting is specific, but the structural problem is universal. Every major AI-assisted information product, from Perplexity to Microsoft Copilot to the retrieval-augmented tools now standard in financial services and policy research, operates on the same three-step logic: retrieve a source, attach a passage, present the result as supported.
For a multilateral institution publishing technical guidance, or a financial services firm whose regulatory positions appear in AI-generated summaries, the implication is direct. The model may cite your document accurately in the narrow sense: the text it quotes will be genuine. The claim it makes using that text may not follow from it at all. Your institution gets the citation credit; the claim attached to your name may be one you would dispute.
This is not a hypothetical risk in the UN system or at the World Bank. Organisations like UNDRR or CGAP publish guidance that is routinely indexed and retrieved by LLMs answering policy questions. If a model cites the Sendai Framework to support a claim the framework does not in fact make, the framework's credibility absorbs the error while the model moves on.