Web search mode alters ChatGPT citation behaviour, audit shows
The model your brand monitors via API is not the model your stakeholders use via chat with web search on.
Key takeaways
- Enabling web search in ChatGPT's chat UI changes citation behaviour systematically, not just occasionally.
- Web-search-enabled responses show lower consistency across repeated runs of the same prompt.
- Most standard AI benchmarks test only the API modality, missing how the chat UI behaves in deployment.
- Brands monitoring LLM visibility with single-run API queries may be profiling a different system than users actually experience.
- Reliable citation audits require repeated runs across both API and chat UI modalities with web search enabled.
Web search mode changes not just what ChatGPT says, but how it justifies what it says. A 4,812-response audit published on arXiv finds that enabling web search in ChatGPT's chat interface produces materially different citation behaviour from the same model accessed via OpenAI's API, a gap that most standard benchmark evaluations never measure because they test only one modality at a time.
The researchers used 401 prompts drawn from two widely used benchmarks, BBQ and SafetyBench, running each prompt three times across both the chat UI and the API, with and without web search enabled. The chat UI with web search produced the most distinctive output profile: higher citation rates, lower response consistency, and different abstention behaviour than the API condition. In plain terms, the same underlying model behaves differently depending on whether web search is on, and that difference is systematic enough to show up across thousands of responses.
The citation architecture most brands are ignoring
The conventional assumption among brand and communications teams is that their LLM visibility problem is a content problem: write better material, earn more citations. That framing misses something structural. If ChatGPT's citation behaviour changes substantially when web search is active, then the question is not just whether your content exists on the web, but whether it surfaces in the retrieval layer that feeds the chat interface at the moment of a query.
Web search mode does not simply add sources to a fixed answer. It reshapes the answer's construction. The arXiv study found that web-search-enabled responses showed lower consistency across repeated runs of the same prompt, which means the retrieved pages are not stable inputs. They vary by run. A brand that appears in the citation layer on one query may be absent from the next, even when the prompt is identical. For a multilateral institution like UNDRR producing authoritative risk data, or a financial services firm whose research is genuinely source-worthy, that inconsistency is not a minor technical footnote. It means citations the institution believes it is earning may be ephemeral.
Modality matters more than model version
The study's sharpest finding is about measurement, not just behaviour. Most AI evaluations, including those used to certify models as safe or deployment-ready, assess only the API modality. The chat UI, which is where the overwhelming majority of non-developer users encounter these models, is effectively untested in standard benchmarks. The implication: the safety and reliability claims made about a given model version may not hold for the version of that model most people actually use.
For enterprise communicators, this has a direct read-across. If your organisation has built an AI visibility strategy based on how the API behaves, because that is what most third-party monitoring tools query, you may be profiling a different system than the one your stakeholders are actually using. ChatGPT's chat UI with web search is the default experience for millions of users asking questions about climate risk, financial regulation, labour market trends, and industrial standards. The API is not.
The response consistency finding sharpens this further. Lower consistency in web-search mode means a brand's citation presence is harder to audit reliably. A single-run scrape of LLM outputs, which is how many brand-monitoring tools operate, will capture a slice of a variable system. A brand that appears cited 60% of the time on a proper repeated-run audit might appear cited 90% or 30% on any single pull, depending on what the retrieval layer returned that day.
The remedy is measurement methodology, not content volume. Brands monitoring their LLM visibility need repeated-run protocols across both API and chat UI modalities, with web search enabled in the chat condition. That is not how most organisations are currently set up. The arXiv study was designed to benchmark AI safety evaluations, but its methodological critique applies with equal force to any entity trying to understand how it is represented in deployed AI systems. Industrial groups publishing technical standards, policy institutions producing governance frameworks, and UN agencies releasing field data all face the same structural blind spot: they are measuring LLM behaviour in a laboratory condition that does not match the field.
The gap between how AI is tested and how it is used is now a brand-visibility problem as much as a safety one.