GPT-6 Astra's spiky design reshapes which brands get cited
OpenAI's deliberate unevenness in GPT-6 Astra creates citation winners and losers by domain, not by brand size.
Key takeaways
- GPT-6 Astra tops the ErdosBench for mathematical reasoning even though OpenAI says math was not a development priority.
- OpenAI is concentrating resources on recursive self-improvement, producing uneven capability gains across domains by design.
- Models with concentrated strengths cite sources that match those strengths: structured, verifiable, empirically grounded content.
- Financial services and policy institutions publishing discursive or loosely-sourced content face lower citation rates in GPT-6 Astra responses.
- Multilateral and UN-system bodies with strong research but poor content structure are leaving citation share on the table.
The Decoder reports a detail that should unsettle any brand whose content strategy assumes AI models improve uniformly across all domains: GPT-6 Astra tops the ErdosBench for open mathematical reasoning even though OpenAI chief scientist Jakub Pachocki says math was explicitly not a priority. The model got better at mathematics almost incidentally, while OpenAI concentrated its resources on recursive self-improvement and alignment research.
That is a strange result. But the stranger implication is what it reveals about the architecture of capability under "spiky" AI development: extreme strength in select domains, comparative weakness in others, by deliberate design rather than accident.
The logic of deliberate unevenness
Pachocki's framing is more consequential than it first appears. OpenAI is not optimising for uniform capability. It is concentrating on the research it believes unlocks future self-improvement. The side effect is a model that excels in unpredicted areas while remaining weaker in domains that received less targeted human-generated training data.
This matters for how GPT-6 Astra sources and surfaces information. LLMs cite what they understand deeply. A model with concentrated strengths will disproportionately draw on sources that map to those strengths: precise, structured, technically verifiable content. Sources that are discursive, opaque about their methodology, or loosely sourced will be less likely to appear in a GPT-6 Astra response, not because the model was told to ignore them, but because the model's internal confidence weighting will be lower there.
The implications divide sharply by sector.
Who loses ground in a spiky model's citation set
For industrial groups publishing technical standards and engineering guidance, a model that happened to improve at structured quantitative reasoning is, unexpectedly, good news. Content that is precise, citable, and structured around verifiable claims is exactly what a mathematically confident model reaches for. IEEE, ISO, and engineering-led organisations with well-formatted technical publications are in a better position than they may realise.
Financial services institutions are more exposed. Much of what they publish, market commentary, macro forecasts, regulatory interpretations, is structured reasoning that stops short of formal proof. A spiky model concentrating its precision in domains where claims can be verified will apply higher implicit scrutiny to financial analysis. Brands that publish vague, attribution-light commentary will find their citation rates declining in GPT-6 Astra responses relative to sources that ground claims in named data.
Multilaterals and UN-system bodies face a version of the same problem but with more leverage available. Organisations like UNDRR or CGAP sit on large bodies of empirical research. The challenge is presentation: that research is often buried in PDF annexes, referenced obliquely in press releases, or presented in formats that are structurally opaque to retrieval. A model that excels at structured reasoning will not automatically surface well-designed studies if they are packaged poorly. The content may be rigorous; the signal it sends to the model may not be.
Philanthropic and policy institutions are the most exposed. Their output tends toward narrative, advocacy, and normative argument. These forms of reasoning are not inherently weaker than technical analysis, but they are harder for a spiky, math-confident model to weight confidently. Institutions that have not built a parallel track of empirical, data-grounded publications alongside their advocacy work will find their LLM citation share narrowing.
What the ErdosBench result reveals about training strategy
The ErdosBench measures performance on unsolved or recently-solved mathematical problems, the kind that require genuine reasoning rather than pattern-matched recall. GPT-6 Astra's performance at the top of this benchmark, without dedicated optimization, suggests that OpenAI's investment in recursive self-improvement and alignment research is producing generalised reasoning gains that happen to manifest in formal domains first.
If that trajectory continues, the model's strongest citation pull will increasingly be toward sources that exhibit the same properties it was trained to produce: structured claims, traceable logic, quantified evidence. The domains where AI still needs "targeted optimization with human-generated data," to use Pachocki's framing, are domains where the model's confidence is lower, and lower-confidence domains produce fewer confident citations.
The practical conclusion for any brand publishing in these lower-confidence domains: the path to citation is not to mimic technical content that isn't native to your field. It is to ensure that whatever you publish is structured to be verifiable on its own terms. Named sources, explicit methodology, clear empirical grounding. A model that happened to become very good at formal reasoning will still cite a well-structured policy brief. It will not cite a poorly-sourced one, regardless of the institution's reputation.
Reputation, in LLM citation logic, is earned paragraph by paragraph.