Pew: one-third of post-2022 web pages show AI-written text
As synthetic text floods .com domains, institutional and multilateral publishers hold a structural citation advantage they must act on now.
Key takeaways
- Over a third of English-language web pages published after ChatGPT's launch show signs of AI-generated text, per Pew Research Center.
- Commercial .com sites are ten times more likely to contain AI text than .edu or .gov domains.
- LLM retrieval systems will apply stronger provenance filters as synthetic content grows; institutionally authored content gains citation advantage.
- Financial services and industrial brands publishing on .com must add verifiable signals: named authors, proprietary data, traceable citations.
- Multilateral and policy institutions risk squandering their credibility advantage by keeping primary research in inaccessible formats.
The Pew Research Center sampled nearly half a million English-language web pages and found that more than a third of those published after ChatGPT's launch in late 2022 show signs of machine-written text. The Decoder reports the finding this week. That single statistic is not a data-quality footnote. It is a structural shift in the information environment that every large language model now ingests, and it changes the logic of brand visibility in AI-generated answers.
The mechanism is worth tracing precisely. LLMs learn from web text; they also cite web text when serving answers in real time via retrieval-augmented generation. If a third of post-2022 pages are themselves AI-generated, the training and retrieval corpus is increasingly a hall of mirrors: models trained on model output, citing pages produced by earlier model runs. The question for any institution trying to appear authoritatively in an LLM's answer is no longer only "does the model know our content?" but "does the model treat our content as a credible signal, or as more noise in an increasingly synthetic corpus?"
The ten-to-one gap that decides who gets cited
Pew's breakdown by domain type is the sharper finding. Commercial .com sites are ten times more likely to contain AI-generated text than .edu or .gov domains. That gap is, for now, an advantage for institutions whose editorial norms preclude bulk AI content generation. Universities, intergovernmental bodies, and public-sector agencies sit in a relatively clean segment of the corpus. A UN agency publishing a rigorous policy brief, a World Bank research group releasing primary data, or a standards body like ISO publishing a normative document occupies rare territory: human-authored, institutionally vouched, and structurally distinct from the .com flood.
LLMs are not indifferent to this distinction. Models trained on quality signals and fine-tuned for factual accuracy develop implicit preferences for authoritative domains. Retrieval systems, whether built into Perplexity, ChatGPT's web browsing, or Google's AI Overviews, apply domain-authority filters that inherited decades of search-era signals. The Pew data suggests those filters are about to do more work, not less, as the .com tier fills with undifferentiated synthetic text.
For multilateral institutions and philanthropic bodies, the implication is that their natural content discipline is becoming a competitive asset in AI retrieval. The risk is failing to recognise it as such, and allowing primary research to remain locked in PDFs, behind registration walls, or in formats that retrieval systems cannot easily parse. Visibility in an LLM answer requires being findable and readable by a crawler, not just authoritative in principle.
For financial services firms and major industrial groups, the picture is less comfortable. These organisations publish heavily on .com domains. Their thought-leadership content, market commentary, and product documentation now competes in the tier that Pew flags as most saturated with AI text. Differentiation requires signals that synthetic content cannot replicate at scale: proprietary data, named human analysts, methodology disclosures, traceable citations. A research note attributed to a named economist with a verifiable track record is harder for a retrieval system to dismiss than an anonymous article optimised for search intent.