Anthropic's Opus 5 reshapes the competitive model landscape
When a cheaper model wins on novel reasoning, enterprise defaults shift and so does which brands get cited in AI outputs.
Key takeaways
- Opus 5 scores 30.2% on ARC-AGI-3, nearly four times higher than GPT-5.6 Sol, at half the token price of Fable 5.
- Citation behaviour varies across models; a source visible in GPT-based outputs has no guaranteed visibility in Opus 5.
- Anthropic targets coding and knowledge work, the high-volume enterprise segments most likely to shift on price.
- Financial services, UN-system bodies, and industrial groups routing AI workflows face an immediate citation-surface risk.
- Brands optimising for a single model's retrieval preferences are repeating the single-search-engine mistake.
The Decoder reports that Anthropic's Claude Opus 5 scores 30.2% on ARC-AGI-3, nearly four times higher than GPT-5.6 Sol, while pricing tokens at roughly half the rate of Fable 5. For a benchmark designed to test novel problem-solving rather than pattern recall, that gap is not incremental. It signals a material shift in which models enterprises will route their highest-stakes queries through.
Benchmark scores are a notoriously unreliable proxy for real-world performance. But ARC-AGI-3 is specifically constructed to resist memorisation, which makes Opus 5's lead harder to dismiss as training-set contamination. When a model can reason through genuinely novel problems at competitive cost, the calculus for enterprise deployment changes: the question is no longer whether to use a frontier model, but which frontier model earns the default routing position.
Why routing position matters more than raw capability
In LLM-powered search and synthesis tools, the model at the top of the default stack sees the most queries. It is also the model that determines which sources get cited. Perplexity, Claude.ai, and an expanding set of enterprise copilots all let users or administrators choose their underlying model, and those choices are increasingly cost-weighted. If Opus 5 delivers near-parity performance to the market's most capable model at half the token price, procurement teams will move toward it. That means Opus 5 becomes the retrieval engine for a larger share of professional queries.
For brands that have built visibility strategies around the citation patterns of GPT-4o or the current Gemini stack, this is the operational problem: citation behaviour is not uniform across models. A source that ranks well in one model's training and retrieval preferences may rank poorly in another's. Anthropic's Claude models have historically favoured long-form, structured content with clear authorial sourcing. If Opus 5 gains share in enterprise deployments, the content signals that earn citations in that model's outputs become more commercially significant.
The sectors most exposed are those that already run large-scale knowledge workflows through LLM interfaces. Financial services firms using AI for regulatory research, multilaterals like the UN agencies and World Bank-affiliated bodies that route policy queries through enterprise AI, and large industrial groups deploying AI copilots for engineering and compliance work all face the same issue: if the model changes, the citation surface changes with it. A technical white paper from a trade association that earned consistent citations in GPT-based outputs has no guarantee of equivalent visibility in an Opus 5-powered workflow.
Anthropic has also positioned Opus 5 explicitly for coding and knowledge work, the two task categories most likely to produce structured, citable output. That positioning is deliberate. Enterprise buyers in those domains are high-volume token consumers, which is exactly the segment that responds to a 50% price reduction. The addressable shift in market share is not trivial.