New AEO tracker benchmarks citation picks across frontier LLMs
Citation behaviour varies by model, not just by content quality. Astra makes those differences measurable for the first time.
Key takeaways
- Citation rates differ significantly across ChatGPT, Gemini, Perplexity and Claude. Ranking well on one model does not guarantee visibility on another.
- Models favour sources that are specific, structured, and already trusted in training data. Broad claims lose to narrow, documented ones.
- Paywalled or poorly indexed research is invisible to the retrieval layer before credibility even becomes a factor.
- Systematic citation measurement is now possible. Brands that adopt it in 2025 will build a compound advantage over those that wait.
- For multilaterals and standards bodies, this tracker reframes the question: not 'are we publishing?' but 'are LLMs citing us, and on which topics?'
Latent Space published its Astra AEO tracker this week, and the timing is pointed. Answer engine optimisation has spent two years generating consultant decks and conference panels; it now has a benchmarking tool that forces the conversation onto empirical ground.
The tracker measures which sources frontier models actually cite when answering queries, across ChatGPT, Gemini, Perplexity, Claude, and others. The operative word is "actually." Most AEO advice has been reverse-engineered from intuition or from the SEO playbook with the labels swapped. Astra generates real prompts, collects real model outputs, and records which domains get picked. That is a different category of evidence.
Citation is not uniform, and the gap matters
The first thing the tracker reveals is that citation behaviour varies meaningfully across models, not just in frequency but in the type of source each model favours. A domain that earns consistent placement in Perplexity's answers may barely register in ChatGPT's. For a brand operating across both surfaces, those are two separate optimisation problems, not one.
This has direct consequences for institutions where authoritative citation is reputationally significant. A multilateral publishing a technical report on climate finance, or a standards body like ISO releasing updated guidance, cannot assume that being well-indexed on Google translates into being cited by frontier LLMs. The models have their own preferences, shaped by training data, retrieval architecture, and reinforcement from user feedback. Astra makes those preferences legible for the first time in a systematic way.
The tracker also surfaces something the SEO analogy obscures: in traditional search, ranking is one-dimensional. Position one beats position two. In LLM citation, the question is whether a source appears in a generated answer at all, and if so, whether it is named explicitly or merely absorbed into the model's output without attribution. Brands accustomed to measuring visibility through rank position need a different instrument. This is that instrument.
What the data says about who wins
Per Latent Space, the Astra tracker was built partly in response to demand from founders and developer-experience leaders who wanted to understand citation dynamics rather than speculate about them. That origin matters because it means the tool is calibrated to queries with commercial and professional stakes, not trivial lookups.
Early patterns from the tracker point toward a familiar but useful finding: models disproportionately cite sources that are already trusted in their training data, that publish structured and retrievable content, and that produce material at the intersection of specificity and authority. A blog post making a broad claim loses to a white paper making a narrow, documented one. A press release loses to a third-party analysis of the same facts.
For financial services firms and philanthropic institutions, the implication is concrete. If your research is paywalled, lightly indexed, or published in formats that resist extraction, you are invisible to the retrieval layer before the model even considers whether your content is credible. Visibility in AI answers begins upstream, at the point where content is structured and distributed, not at the point where the prompt is written.
The benchmark as a strategic lever
Benchmarking tools create accountability. Once citation rates are measurable, the question of whether a communications or content investment is working has a quantifiable answer. That is new. For senior marketers at industrial groups or UN-system bodies, it changes what the brief to an agency looks like. The question shifts from "are we producing thought leadership?" to "are frontier models citing our thought leadership, and which ones, and on which topics?"
Astra does not yet cover every model or every query type, and any tracker built on prompt sampling carries the usual caveats about representativeness. But the methodological direction is right. The alternative, which is to optimise for AI citation based on vendor documentation and anecdote, is not a strategy. It is a guess dressed in confidence.
The deeper signal here is competitive. Brands that adopt systematic citation measurement in 2025 will accumulate a compound advantage over those that wait. Citation data reveals not just where you stand but where your peers stand and which content categories are underserved by credible sources. That is a content opportunity with a traceable return, which is a rarer thing than it sounds.