Why one citation can lock in a brand across a chat
Winning the first citation in a chatbot conversation may matter more than winning any single answer later on.
Key takeaways
- A source cited in turn one of a chatbot conversation becomes more likely to be cited in later turns.
- The effect, called conversational capture, runs through both retrieval history and user follow-up phrasing.
- Single-answer GEO audits miss this compounding dynamic entirely.
- Institutions with existing authority on a topic stand to benefit most from first-turn capture.
- Winning the opening question in a topic area may matter more than optimising for long-tail follow-ups.
A source cited in the first turn of a chatbot conversation becomes substantially more likely to be cited again, according to a new framework from researchers studying generative engine optimisation. The paper, "Conversational Capture: A Trajectory-Level Framework for Evaluating Generative Engine Optimization in Multi-turn Human-Agent Interaction," published on arXiv, formalises something most GEO vendors have been pricing as a feature without naming it: early citation compounds.
The researchers call it conversational capture, and the mechanism has two moving parts. The first is machine-side: retrieval-augmented systems condition on conversation history, so a source that entered the context window in turn one has a structural advantage in turn three, independent of relevance. The second is human-side: a user who reads a brand name in an early answer tends to ask follow-up questions that reference it, which steers retrieval right back to that source. The loop is self-reinforcing. The paper formalises this as a two-layer closed system and proposes trajectory-level metrics, including cumulative conversational visibility, to replace the single-answer snapshot that has defined GEO measurement until now.
The metric everyone has been using is the wrong one
Almost every visibility audit sold to marketing teams today answers one question: for a given prompt, does ChatGPT, Perplexity or Gemini cite your brand? That is a single frame from a film. The arXiv paper's contribution is to argue, with a formal model rather than a hunch, that the film matters more than the frame. If capture is real and measurable, then a source's odds of appearing in turn five are not independent of what happened in turn one. They are conditional on it. Treating each answer as a discrete event, the way most GEO dashboards do, throws away the one variable that predicts long-run visibility: who got there first.
This has an uncomfortable implication for brands that are currently losing on single-answer visibility audits and assuming they can compete answer by answer. If capture dynamics hold at scale, the real contest is not for share of citations across many independent queries. It is for the first citation in conversations that are likely to run long, because that citation seeds a multi-turn asymmetry that compounds. A financial services brand cited once in response to "what are the capital requirements under Basel III" may find itself cited again when the user asks three follow-ups about implementation timelines, simply because the model's context window now contains that brand's name and the user's own language has started to mirror it.
Who this favours, and who it leaves exposed
The mechanism rewards incumbency inside a conversation in a way that single-turn GEO metrics cannot see. For multilateral institutions and standards bodies, this cuts two ways. A body like ISO or a UN agency that gets cited first on a technical or regulatory question stands to benefit from a kind of conversational incumbency that mirrors their offline authority: once the model has anchored on the UN Framework Convention on Climate Change as a source, follow-up questions about compliance or reporting are more likely to pull from the same well. That is good news for institutions whose authority is already encoded in training data and retrieval corpora.
It is less comfortable reading for challenger brands, newer entrants, or any organisation whose content has not yet been indexed deeply enough to win that first citation. Under a capture dynamic, losing the opening exchange is not a single lost impression. It is a compounding disadvantage across the rest of the session, because the model's own history-conditioning and the user's own follow-up phrasing both work against the challenger from that point forward. A philanthropic foundation competing for visibility against a government agency on a policy topic faces a steeper hill than single-answer audits suggest, because the agency's first-mover citation seeds every subsequent turn.
The practical upshot for anyone measuring AI visibility is that answer-level tracking, the current industry default, is measuring the wrong unit. A brand that shows up in 15% of single-turn test prompts but never captures the opening turn in realistic multi-turn sessions may be invisible where it counts. Content strategy built to win first-turn citations, meaning content structured to answer the most likely opening question in a topic area rather than the long tail of follow-ups, becomes disproportionately valuable. GEO vendors who keep selling single-answer scorecards are selling a photograph of a race that is actually a relay.