ChatGPT picks markets by query language, not exit IP
Language corpus, not exit IP, controls LLM brand selection. English-language content gaps cost B2B brands visibility in every market.
Key takeaways
- Query language predicts which brands ChatGPT recommends; exit IP has no measurable effect.
- Local suppliers won 24 of 24 local-language runs; English queries from the same IPs returned global brands 6 of 6 times.
- Top recommendations changed across identical runs on 4 of 6 prompts, making single-session spot-checks unreliable.
- Multilingual content investment is a prerequisite for market-level AI visibility, not an optional localisation task.
- Brands must audit English-language corpus depth first, then assess which local-language corpora matter for each target market.
ChatGPT answered a commercial query in Estonian from an Estonian IP address and named a local supplier. Asked the same question in English from the same connection, it named a global brand instead. The query language changed; the exit IP did not. The recommendation changed entirely.
That finding, reported in a controlled probe published on arXiv, has direct consequences for any brand competing in multilingual markets. The researchers ran 234 sessions against the logged-out ChatGPT web interface and the OpenAI API across four exit countries and six query languages, with six identical runs per cell, on 29 and 30 August 2026. The results are worth reading carefully because they dismantle two assumptions that most marketing teams have quietly built into their AI-visibility strategies.
The IP assumption is wrong
The dominant intuition in B2B marketing circles is that LLM-powered search behaves roughly like Google local: the model detects your location and serves regionally relevant results. That is how teams have been thinking about AI visibility, and it is incorrect. In this study, query language predicted market selection; exit IP did not. Asked in English on connections routed through Estonia and Turkiye, ChatGPT named global brands in both cases, zero for six. Asked in the local language on the same connections, local suppliers appeared twenty-four times in twenty-four runs. Language and location are separable factors, and language wins.
The mechanism behind this is not mysterious. Language training data is not distributed evenly across geographies. A model trained predominantly on English-language text has learned which brands English speakers associate with a given category. When you query in English, you invoke that corpus. When you query in Estonian, you invoke a narrower, locally anchored corpus, and the brands in it are different. The model is not doing geography; it is doing pattern-matching against language-specific training distributions.
For a multinational industrial group or a financial institution operating across the EU's smaller markets, that distinction is consequential. If the buyers, procurement leads, or policy researchers who matter most to your pipeline tend to query in English, even when they are based in Riga or Tallinn, the model will consistently surface global competitors. Local brand investment, however substantial, does not compensate for an English-language content gap.
Instability compounds the problem
The second finding is just as uncomfortable. The top recommendation changed across six identical runs on four of six prompts. That rate was the same in the browser interface and in the API, with web search both enabled and disabled. The researchers conclude that instability is a property of the system, not of the surface. There is no stable rank-one position in ChatGPT's commercial recommendations. A brand appearing first in one session may disappear entirely in the next.
This matters for measurement as much as for strategy. Teams that have been spot-checking their ChatGPT visibility, running a handful of queries and reporting back a position, are measuring noise. The signal is the distribution of appearances across many runs; any single run tells you almost nothing. For multilaterals like UNDRR or institutions in the UN system that rely on AI-generated summaries to shape procurement research, this variability means that even a well-positioned brand may not appear when it counts.