Ahrefs data upends common AI search assumptions
Backlink authority does not predict AI citation. Ahrefs' data points to topical depth and third-party corpus presence as the real drivers.
Key takeaways
- Domain authority does not reliably predict citation in AI search results.
- Topical specificity beats broad, high-authority pages when queries are narrow and precise.
- Structured data and schema markup show minimal correlation with AI citation rates.
- AI models compound advantage for brands with wide third-party editorial coverage.
- Earned media and external citations are the primary lever for LLM visibility, not owned content.
Ahrefs analysed 15 million data points to test what actually drives citation in AI search. Search Engine Journal published the findings, and several of them contradict advice that has become received wisdom among SEO practitioners in the past eighteen months.
The headline result: domain authority, that long-standing proxy for credibility, does not reliably predict whether a page gets cited by AI models. Pages with relatively modest authority scores appear in AI answers at rates comparable to high-authority domains when their content directly matches query intent. For brands that have spent years accumulating backlinks as a citation strategy, this is an uncomfortable finding. The mechanism that made Google tick does not map cleanly onto the retrieval logic of large language models.
What the data actually shows
The Ahrefs study found that AI models consistently favour topical specificity over domain-level signals. A page that answers a narrow question precisely outperforms a broader, higher-authority page that touches the topic in passing. This matters most for organisations that publish content at scale without a clear topical architecture: think large multilateral institutions with sprawling resource libraries, or financial services firms whose content teams produce a steady stream of thought leadership that covers many subjects at shallow depth.
Freshness, too, is more nuanced than the conventional wisdom holds. Models do not simply prefer newer content. They prefer content that was current at the time a question is most frequently asked, which is a subtler distinction. Evergreen content that has been methodically updated outperforms both stale older pages and brand-new content with thin history. For policy institutions and UN-system bodies publishing annual reports or multi-year research, the implication is that structured, versioned updating of existing pages is more valuable than creating new ones.
The study also challenged the assumption that structured data and schema markup significantly improve AI citation rates. Ahrefs found minimal correlation. The models appear to extract meaning directly from prose, not from markup tags. This does not mean structured data is worthless for traditional search; it means brands should not expect it to serve as a shortcut into AI answers.
The citation gap that compound interest creates
Perhaps the sharpest finding concerns brand familiarity. Entities that AI models encounter repeatedly across their training and retrieval corpus get cited more often, independent of whether any individual page is technically superior. This is a compounding effect. Brands already well-represented across trusted third-party sources, Wikipedia, industry publications, standards bodies, grow their citation probability with each additional mention. Brands that have relied on owned channels and paid media, without investing in genuine third-party editorial coverage, find themselves systematically underweighted.
For industrial groups and major procurement-driven enterprises, this matters at precisely the moment it is hardest to fix. When a procurement officer at a large infrastructure firm uses an AI assistant to shortlist vendors, the model draws on a body of third-party evidence. A brand with deep technical capability but thin external coverage loses ground to a competitor with broader mentions, even if the competitor's content is less rigorous. The authority that counts is not domain authority; it is corpus presence.
Philanthropic and policy institutions face a variant of the same problem. Their credibility rests on research quality, but if that research is self-published and rarely cited by external outlets, models may not treat them as authoritative sources on their own areas of expertise. Publishing a rigorous report is necessary but not sufficient. The report must enter the broader information ecosystem to register as a citation signal.
The practical conclusion from Ahrefs' 15 million data points is blunt: brands optimising for AI visibility need to treat earned media and third-party coverage as the primary investment, not a supplement to owned content. Domain authority built through backlinks was a proxy for trust in a link-graph world. In a retrieval-augmented world, the proxy is corpus breadth, and the only way to build it is to produce content that other credible sources actually reference.