Opus 5 leaps ahead on the benchmark built to test real reasoning
A 4x benchmark gap on a test designed to resist memorisation tells enterprise buyers something about which model will cite their content accurately.
Key takeaways
- Claude Opus 5 scored 30.2% on ARC-AGI-3, nearly four times GPT-5.6 Sol's record of 7.8%.
- ARC-AGI-3 is designed to test genuine generalisation, not pattern recall, making this gap qualitatively significant.
- Opus 5 independently formulated reflection equations, a behaviour the benchmark's developers had never seen from another model.
- Brands targeting complex, analytical queries are better positioned when the citing model can handle genuine abstraction.
- Enterprise adoption of Opus 5 in knowledge and research workflows would shift citation patterns toward Claude-based outputs.
The ARC-AGI-3 benchmark was designed to be hard by people who found previous benchmarks too easy. Claude Opus 5 just made it look less hard. The Decoder reports that Anthropic's new model scored 30.2% on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8%. Fable 5, the other named competitor, did not come close either.
That gap is not a rounding error. It is a structural claim about the kind of reasoning Opus 5 can do that its peers presently cannot.
What ARC-AGI-3 actually tests
ARC-AGI-3 is the third iteration of François Chollet's Abstraction and Reasoning Corpus, a benchmark explicitly designed to resist pattern-matching. Earlier versions tested novel visual analogies; this iteration is harder still, built to identify generalisation that cannot be achieved by memorising training data. A model that scores well on it has, in some meaningful sense, done something new, not recalled something familiar.
The benchmark's developers noted that Opus 5 independently formulated "reflection equations," a behaviour they had not previously observed in any other model. That single observation is more interesting than the headline number. A model behaving in ways its evaluators had not anticipated is the clearest signal yet that the capability frontier is moving in a direction that quantitative scores alone cannot fully describe.
GPT-5.6 Sol managed 7.8%. Human performance on ARC-AGI-3 sits in the 85% range, so no one should mistake 30.2% for a solved problem. But the distance between 7.8% and 30.2% on a task designed to prevent incremental gains is substantial. It implies that Opus 5 is doing something qualitatively different, not simply applying more compute to the same approach.
What this means for which model ends up in AI answers
Model capability and AI-search citation share are not the same thing. ChatGPT's share of AI-generated citations in enterprise research queries remains large because OpenAI has distribution advantages baked into Microsoft's infrastructure, from Azure to Copilot to the default integrations many large organisations already have in place. A benchmark lead does not automatically translate into a citation lead.
But the relationship between underlying reasoning quality and citation reliability is tightening. For senior communicators at multilateral institutions, industrial groups, or financial services firms, the relevant question is not which model scores higher on an abstract test. It is which model is being used by the people and systems making consequential decisions about information retrieval. As organisations evaluate AI deployments for internal knowledge management, policy analysis, and market intelligence, Opus 5's demonstrated capacity for novel logical inference becomes a procurement argument, not just a technical footnote.
Anthropic has historically been stronger in the research and professional services segments than in consumer-facing deployments. A benchmark result of this kind reinforces that positioning and gives enterprise buyers a legible signal to point to when justifying a move toward Claude-based tooling. If those buyers move, the citation patterns they generate will follow.