OpenAI launches GPT-6 Astra to rival Claude Fable
A more capable frontier model rewards specific, structured content and penalises vague positioning — visibility built for weaker models may not survive the upgrade.
Key takeaways
- GPT-6 Astra scores 99.9% on ARC-AGI 3, though the $19K compute cost and custom configuration limit direct comparisons with Fable 5.
- OpenAI priced Astra identically to Claude Fable 5 and 5.1 at $10/million input and $50/million output tokens, signalling a stable frontier price tier.
- Astra rolls out to ChatGPT Plus, Pro, Business, and Enterprise users within days, meaning citation behaviour inside ChatGPT shifts rapidly.
- Stronger reasoning models are more selective about sources; brands with thin or fragmented content face measurable citation loss.
- Multilaterals and industrial groups with dense, inconsistently indexed publications are most exposed to stricter model coherence standards.
OpenAI released GPT-6 Astra on 3 September, and the headline number is striking: 99.9% on the ARC-AGI 3 benchmark, a reasoning test released only in March 2026. Simon Willison's Weblog flags the caveat immediately. That score cost $19,000 in compute using a custom "Provider Adapter" configuration, and Anthropic has not yet published a comparable Fable 5 result. The benchmark is real; the comparison is, for now, one-sided.
The more durable signal is the pricing decision. OpenAI set Astra at $10 per million input tokens and $50 per million output tokens. Those are exactly the rates for Claude Fable 5 and 5.1. When a market leader matches a competitor's price to the dollar rather than undercutting it, the message is that both firms believe this tier can hold. Frontier reasoning is becoming a commodity at a specific, visible price point.
What changes inside the answer engine
Astra is not yet widely available. At launch it is rolling out to a limited set of organisations, with broader access to ChatGPT Plus, Pro, Business, and Enterprise users following in the coming days, alongside API and AWS availability. That rollout sequence matters for enterprise buyers, because the models powering ChatGPT's responses will shift as Astra propagates.
For brands whose visibility depends on what ChatGPT cites and synthesises, a more capable reasoning model cuts both ways. Better reasoning means the model is harder to satisfy with thin or repetitive content; it is more likely to identify the authoritative source among several candidates and cite that one rather than averaging across them. A financial services firm whose thought leadership currently appears in ChatGPT responses because it ranks well in a simpler retrieval step may find a more discriminating model weights structured, citable evidence more heavily.
Multilateral institutions and UN-system bodies face a specific version of this risk. Their publications are authoritative but often dense, formatted for internal audiences, and inconsistently indexed. A frontier model with stronger reasoning can parse complex documents more accurately, but it also applies stricter coherence standards. A policy brief that loses its argument in procedural language will not be retrieved reliably simply because the institution's name carries weight.
The industrial and philanthropic sectors have a parallel concern. When GPT-6 Astra is processing a query about, say, cement decarbonisation or financial inclusion in low-income markets, it will weigh the clearest, most self-consistent source. Being the largest voice in a sector is no longer sufficient if that voice is distributed across a dozen disconnected documents with no consistent framing.
The Fable comparison is unfinished
Willison is precise about what is not yet known. He has not tested Astra himself. Anthropic has not published an ARC-AGI 3 score for Fable 5. The benchmarks OpenAI cites are self-reported. This is not unusual at launch, but it means the competitive picture is incomplete, and enterprise buyers should treat capability claims cautiously until independent evaluations accumulate.