Study: Chinese AI search cites brands at 8.3% selection rate
Being retrieved into an AI citation pool is not the same as appearing in the answer. For most brands, synthesis is where visibility ends.
Key takeaways
- Chinese AI search platforms surface only 8.3% of brands in their citation pools inside the final generated answer.
- 88% of sources containing contact information were retrieved but stripped before the answer reached the user.
- Cross-source corroboration is a key predictor of surfacing: one-location content is structurally disadvantaged.
- The two-stage retrieve-then-synthesise architecture behind this finding applies to all major generative search systems, not just Chinese platforms.
- B2B brands that publish authority content in a single channel face a compounding visibility loss at the synthesis stage.
Eight point three percent. That is the share of brands appearing in Chinese generative search citation pools that actually surface in the final answer, according to a large-scale empirical study published on arXiv examining four mainstream Chinese-language AI search platforms across 614 queries and more than 160,000 citation-level records.
The number deserves attention not because it is surprising but because it is precise. Until now, most discussion of brand visibility in generative search has rested on qualitative observation or small-sample testing. This study gives practitioners something rarer: a statistically grounded selection rate, derived from a controlled design that replicated each query across eight platform interfaces three times over.
Citation and surfacing are two separate problems
The study's most useful conceptual contribution is the distinction it draws between retrieval and presentation. A source can be retrieved into a platform's citation pool and still never appear in the generated answer. For brands, this creates two failure modes, not one. The first is familiar: failing to be cited at all. The second is subtler and more insidious: being cited in the model's working pool but edited out during answer synthesis.
Twelve point four percent of retrieved sources that contained contact information contributed that contact information to answers. Read that the other way: roughly 88% of sources carrying contact details were retrieved and then discarded before the answer reached the user. For any organisation whose commercial value depends on being findable, that gap between retrieval and surfacing is where visibility dies.
The study identifies three factors that predict whether a retrieved source makes it into the final answer: content fit, cross-source occurrence count, and semantic role. In plain terms, sources that matched the query's intent closely, appeared across multiple independent sources, and played a structurally important role in the answer's argument were more likely to survive the cut. Novelty alone did not help. Presence in one place did not help. The mechanism rewards corroboration.
What this means for brands outside China
The study is confined to Chinese-language platforms, which limits direct extrapolation. The four platforms studied, presumably including the dominant generative search interfaces available in China, operate under regulatory and content-moderation conditions that differ from Western equivalents. Precise platform names are not disclosed in the available abstract.
The underlying mechanism, however, is not culturally specific. Every major generative search system now operates the same two-stage architecture: retrieve first, synthesise second. The selection pressures the study documents, favouring corroborated, intent-matched, semantically central sources, are structural features of retrieval-augmented generation, not quirks of Chinese deployment. An 8.3% selection rate in China is a calibration point, not a ceiling unique to one market.