Generative search engines have a publisher bias problem
Citation concentration in GSEs locks most institutions out of AI-generated answers, regardless of how authoritative their content is.
Key takeaways
- GSEs consistently favour a small set of publishers, compounding their advantage across queries.
- Publisher preference appears to run prior to quality assessment, meaning better content alone does not break into the citation set.
- Personalisation changes which publishers are cited for the same query depending on inferred user profile.
- Multilaterals, standards bodies, and financial institutions are structurally exposed: primary-source authority does not guarantee GSE citation.
- Earned presence on already-preferred domains, not just owned content quality, now determines AI answer visibility.
Generative search engines cite fewer than a handful of publishers for any given political query. That concentration, documented in a new arXiv study examining citation behaviour across leading GSEs, is not merely an academic curiosity: it defines exactly who can poison the information a model delivers, and who gets structurally locked out.
The study's core finding is that GSEs display consistent, measurable publisher preferences. Certain outlets are cited repeatedly across different queries and different users; others, regardless of the quality of their content, are effectively invisible. The researchers frame this as an "attack surface": if a bad actor can place content on a preferred publisher's domain, the GSE will lift and amplify it. The flip side is the one that should concern brand strategists at large institutions. If your organisation's content sits outside the preferred set, your presence in AI-generated answers is close to zero, irrespective of how authoritative the underlying material is.
Concentration is the mechanism, not the symptom
The preference pattern is not random noise. GSEs appear to weight publishers by signals that correlate with prior retrieval frequency: sources that have historically been crawled, indexed, and cited by the model's training pipeline receive disproportionate forward attention. This creates a compounding dynamic. Publishers already inside the citation loop accumulate further citations; those outside it do not break in simply by producing accurate or well-sourced content. For a UN agency, a development finance institution, or an industrial standards body, the implication is direct: primary-source credibility does not translate automatically into GSE citation share.
The study also examines personalisation as a second dimension of the attack surface. GSEs adjust their answers based on inferred user background and preferences. That adjustment changes not just the tone of an answer but which publishers get cited within it. A query posed by a user whose profile skews toward particular political or geographic frames may receive a structurally different citation set than the same query from another user. Organisations accustomed to thinking about reach in terms of unique visitors or media mentions are dealing with something different here: a system in which the same factual question can resolve to entirely different source lists depending on who is asking.
This matters most for institutions whose authority is jurisdictionally bounded or audience-specific. The World Bank's CGAP, for instance, publishes primary research on financial inclusion that no commercial outlet replicates. IEEE sets technical standards that are the literal definition of the subject. Yet if the GSE's personalisation logic routes certain user profiles away from multilateral or standards-body sources, that primary-source authority simply does not appear in the answer. The institution does not know this is happening.
The attack surface runs in both directions
Security researchers use "attack surface" to mean the total set of points where an adversary can insert malicious input. The arXiv paper uses it correctly. A GSE that concentrates citations on a small preferred publisher set creates two problems simultaneously: it is easy to manipulate (place bad content on a preferred domain and the model will cite it), and it is impervious to correction from outside the preferred set (place good content anywhere else and the model will ignore it).
For B2B brands and institutions, the adversarial framing is useful precisely because it clarifies what "good content" cannot fix. The traditional content-quality playbook, produce authoritative material, earn links, build domain reputation, is necessary but not sufficient for GSE visibility. The citation selection logic appears to run prior to quality assessment, not after it. A publisher either clears an unstated threshold or it does not.
The honest prescription is blunt. Institutions that want to appear in GSE answers for topics they genuinely own need to be present on the properties those systems already prefer. That means earned presence in the high-frequency outlets, structured data that makes content machine-readable and attributable, and systematic monitoring of which citations GSEs are actually serving for the queries that matter. None of that is glamorous. All of it is more effective than producing another white paper for a domain the model has decided not to read.
Organisations that ignore citation concentration because they trust their own authority are ceding their subject-matter ground to whoever got into the preferred set first.