GPT-5.6 quietly changed how ChatGPT selects sources
If GPT-5.6 queries specific domains directly, content quality alone no longer determines which brands get cited.
Key takeaways
- The share of ChatGPT Search fanout queries using the site: operator rose sharply after GPT-5.6 rolled out in August 2026.
- Site:-targeted retrieval means GPT-5.6 may pre-select source domains before composing answers, bypassing open-web ranking signals.
- Brands absent from the model's implicit domain shortlist will not be discovered, regardless of content quality.
- Institutions publishing open-access, domain-canonical content at scale hold a structural citation advantage that cannot be built quickly.
- Content hosted on third-party or partner domains rather than an institution's canonical URL is at elevated risk of being missed.
Promptwatch recorded a near-zero baseline turning into a structural feature. For weeks, the share of ChatGPT Search "fanout" queries containing the site: operator hovered between 0.3% and 0.5%. After GPT-5.6 rolled out earlier this month, that share climbed sharply, according to data published by Promptwatch and flagged by Simon Willison's Weblog. The movement is not noise. It is a change in retrieval architecture.
Fanout queries are the sub-searches ChatGPT generates internally to gather sources before composing a reply. When the model appends site:example.com to those queries, it is not browsing the open web and hoping for the best. It is targeting specific domains by name. That is a materially different selection logic, and it has direct consequences for which brands appear in answers and which do not.
What the site: operator actually does to citation selection
In conventional web search, a query without a site operator lets ranking signals determine who surfaces. Domain authority, backlink profiles, freshness, and relevance all compete. The model can pull from anywhere the index surfaces.
A site:-prefixed fanout query bypasses much of that competition. GPT-5.6, if Promptwatch's data reflects real behaviour, has apparently learned to generate queries of the form site:who.int pandemic preparedness or site:imf.org sovereign debt outlook rather than leaving retrieval open. The model is, in effect, pre-selecting its source pool before the answer is composed.
For brands outside that pre-selected pool, the implication is pointed: producing excellent content is necessary but no longer sufficient. If GPT-5.6 has developed a prior belief about which domains answer which question types, a brand that is not on that internal shortlist will not be retrieved regardless of content quality. The model will not discover it. It will query around it.
The mechanism likely reflects how the underlying model was trained. GPT-5.6 has presumably internalised associations between topics and authoritative domains from its training data, and now externalises those associations as explicit site: operators during retrieval. That is more efficient, more predictable, and more opaque to anyone trying to influence the citation output.
Who holds the structural advantage
Institutions with high training-data density on specific topics hold the largest advantage. The World Health Organisation on disease surveillance. The IMF on fiscal policy. IEEE on engineering standards. ISO on certification. These bodies publish prolifically, are cited across the open web, and have presumably trained GPT-5.6 to associate their domains with their respective domains of expertise. They will be queried directly. They will be cited.
For industrial groups, multilaterals, and financial institutions, this is the competitive dynamic worth tracking. A bank with deep public content on climate finance may find its research cited routinely. A peer institution that publishes the same quality of analysis but behind a login wall, or with thin public-facing content, will not be in GPT-5.6's shortlist. Domain-level trust now operates at retrieval time, not just at ranking time.