New EHQ benchmark tests whether LLMs know what they don't know
Which LLMs confabulate about your brand, and which admit they don't know? New benchmark data forces the question.
Key takeaways
- 14 LLMs show substantial variation in epistemic honesty, meaning confabulation risk is model-specific, not uniform.
- Brands underrepresented in training data face the same LLM risk as fabricated entities that don't exist at all.
- A high-restraint model that abstains is safer for brands than a low-restraint model that generates confident wrong answers.
- Post-cutoff events and recent rebrands sit in the same epistemic blind spot for models that don't acknowledge their knowledge limits.
- Brand visibility strategy must account for model-specific confabulation rates, not just whether training data exists.
Fourteen LLMs sat a test most brands have never thought to run. The question was not whether the models are accurate, but whether they admit when they are not. Per arXiv, researchers have published the Epistemic Honesty Quotient (EHQ), a behavioural benchmark applied across a frozen registry of 21 model API routes, of which 14 completed the full confirmatory analysis. The results show "substantial variation" across models on a metric that turns out to matter enormously for brand and entity visibility in AI-generated answers.
The benchmark's architecture is worth understanding precisely because it maps directly onto the failure modes that damage brands. EHQ-3000 comprises 3,000 questions across four category types: Fabricated Entity questions (asking about things that do not exist), Post-Cutoff Event questions (asking about events after a model's training data ends), Hyper-Niche True questions (asking about things that exist but are obscure enough to trip a model into confabulation), and Context-Conditioned questions (asking about entities whose answer depends on context the model may not have). Each category probes a different mechanism of confabulation. A model that scores well restrains itself from answering things it cannot know; one that scores poorly fills the gap with confident-sounding fabrication.
The fabricated-entity problem is not hypothetical
For senior marketers at large institutions, the Fabricated Entity category is the one that should attract immediate attention. An LLM that generates fluent, confident answers about entities that do not exist will, by the same mechanism, generate fluent, confident answers about entities that do exist but are mischaracterised. The model does not distinguish between "I am inventing this" and "I am misremembering this." From the outside, the outputs look identical: authoritative prose, plausible detail, no hedging.
This matters acutely for organisations like multilateral institutions, industrial groups with complex subsidiary structures, or financial services firms with precise regulatory identities. HOLCIM and Adecco, for instance, operate under brand architectures that have changed materially in recent years. A model trained before a rebrand, or trained on sparse data about a specific entity, is structurally identical to a model trained on a fabricated entity. The EHQ Fabricated Entity score is, in effect, a measure of how likely a given model is to confabulate rather than abstain when it lacks reliable knowledge. Brands that are underrepresented in training data face the same epistemic gap as brands that were invented for the purpose of this test.
What the variation across 14 models means in practice
The study's finding of "substantial variation" across models is where the practical implication crystallises. B2B brands and communications teams typically think about AI visibility as a single channel: "what does ChatGPT say about us?" The EHQ results suggest the better framing is: "which models confabulate about us, and which ones abstain?" Those are very different risk profiles.