ChatGPT cites small, unlicensed sites as often as big ones
Content licensing deals with OpenAI may not drive citation priority. Answer-shaped content does.
Key takeaways
- ChatGPT's search index shows no measurable retrieval preference for licensed publishers over unlicensed sites.
- Content licensing deals with OpenAI appear to affect training data, not live citation priority.
- AI citation frequency correlates with answer clarity and structural retrievability, not domain size.
- Large institutions face citation competition from smaller, more direct sources, regardless of brand authority.
- Organisations that format content as direct answers gain visibility that institutional prestige alone cannot deliver.
The conventional assumption in enterprise content strategy is that ChatGPT favours the sources it has paid for. Search Engine Journal reports otherwise: new data finds no statistically meaningful difference in how OpenAI's in-house search index treats licensed publishers versus unlicensed ones. Size, it turns out, confers no structural advantage.
This matters more than it first appears, and not because it is reassuring to small sites.
The mechanism behind the parity
OpenAI's search index operates differently from a traditional web crawler ranking signals. When ChatGPT retrieves live web content to answer a query, it appears to select sources based on topical relevance and retrieval quality, not on the provenance of a content licensing agreement. The model does not, in effect, keep a preferred-supplier list at retrieval time.
That finding cuts against a widely held belief inside large media and publishing groups, several of which have signed content deals with OpenAI on the premise that partnership confers citation priority. If the data holds, those agreements may improve training data coverage but do not appear to translate into retrieval preference when the model answers a live query. The commercial logic of those deals deserves re-examination.
What actually drives citation
Parity between licensed and unlicensed sites implies that ChatGPT's index is optimising for something else entirely. The most plausible candidates are structural: whether a page directly and completely answers the query, whether the prose is unambiguous enough for a model to extract a clean passage, and whether the content exists at a URL the crawler has reached recently enough to trust.
This is consistent with what practitioners already observe in AI Overviews and Perplexity: citation frequency correlates with answer-shaped content, not with domain authority in the traditional SEO sense. A 400-word explainer on a niche policy site can outperform a paywalled Reuters longform if it gives the model a cleaner extraction target.
For financial services firms, multilateral institutions, and major industrial groups, this reframes the competitive problem. The threat to their AI visibility is not that OpenAI has signed deals with their media rivals. The threat is that a specialist NGO, an academic working group, or a well-structured trade blog answers a query more directly than their own communications do. Brand size provides no buffer.
Who wins from a level index
Unlicensed specialists win, provided they produce content that retrieves cleanly. A World Bank working paper formatted for human readers behind a clunky PDF portal loses to a clearly structured HTML summary on a lesser-known development finance site. An ISO standard explained in bureaucratic prose loses to a three-paragraph breakdown on a compliance consultancy's blog.
The implication for large organisations is structural. Producing authoritative long-form content remains necessary for credibility; it is not sufficient for AI citation. The index appears indifferent to institutional prestige. It rewards the page that resolves the query with the least inferential friction.
Licensed publishers that banked on preferential treatment now face a harder editorial question: if the deal does not move the retrieval needle, what does? The answer is the same one available to any site without a deal. Write directly. Structure clearly. Publish at a URL the crawler can reach.
For communications teams at multilaterals and large industrials who have spent the past two years watching media licensing negotiations from the sidelines, this data is quietly good news, with a sharp condition attached: the field is level only for organisations willing to produce content that functions as an answer, not merely as a document.