GSEs erode the web corpus that feeds their own answers
The crawlable web is a common-pool resource, and competing GSEs are overfishing it. Here is what that means for brands whose authority depends on being cited.
Key takeaways
- Competing GSEs face a collective action failure: the symmetric equilibrium extraction rate rises with the number of players, not falls.
- The crawlable corpus degrades on three dimensions at once: volume, quality, and content longevity.
- A single long-run-oriented GSE could self-regulate; a competitive market cannot coordinate that restraint.
- Institutions producing open, crawlable, authoritative content may see relative citation share rise as the broader corpus thins.
- Regulatory or licensing frameworks are the only structural fixes the model identifies; neither is imminent.
Extraction without return is, in economic terms, a commons problem. The arXiv paper "When Search Eats the Web," by researchers modelling the crawlable web as a shared resource, makes that argument with formal rigour. The finding is blunt: generative search engines (GSEs) that answer queries directly from crawled content, without returning traffic to the originating publisher, degrade the very corpus that makes their answers worth reading. Volume shrinks, average quality falls, and content becomes more perishable. At the limit, the corpus goes extinct.
The mechanism is straightforward. Publishers earn revenue from visits. GSEs intercept those visits. Deprived of income, publishers either block crawlers or stop investing in fresh, high-quality content. The paper calls this "extraction without return" and treats the crawlable web as a common-pool resource, analogous to a fishery. One player overfishing does not immediately destroy the stock; but in a multi-player game, the symmetric equilibrium extraction rate is non-decreasing in the number of competitors. More GSEs means more extraction, not a race toward restraint.
The competitive trap
That last result deserves attention. A single long-run-oriented GSE can, the model shows, choose to stay below the erosion threshold. It internalises the cost of corpus degradation because it depends on the corpus for future answers. A market with several competing engines cannot coordinate on that restraint. Each player's dominant strategy is to extract more, not less, because any traffic it leaves for publishers will fund content that competitors can also crawl. The result is a collective action failure with a structurally inevitable direction: toward the threshold.
Google, Perplexity, ChatGPT's search mode, and Microsoft Copilot are already competing in this space. None has a credible mechanism to coordinate extraction rates. Antitrust law, designed to prevent coordination on price, actively discourages the one form of coordination that might preserve the commons.
What this means for brands that produce substantive content
For B2B organisations whose authority rests on published research, policy analysis, or technical documentation, the paper's implications run in two directions at once. In the short run, GSEs cite high-quality, authoritative sources more frequently than thin content. Institutions such as those in the UN system, the World Bank Group, or ISO have a structural advantage: their output is exactly the kind of material that models treat as credible. Perplexity and ChatGPT already surface CGAP financial-inclusion data and UNDRR risk-reduction frameworks in direct answers. Citation is visibility.
In the medium run, however, the erosion dynamic threatens that advantage. If the funding model for quality content collapses across the broader web, models trained or retrieval-augmented on degraded corpora will produce worse answers. The authoritative institutional sources will still be there, but the contextual web that situates and corroborates them will thin out. A citation in a poorly constructed answer is not the same as a citation in a well-constructed one.
Large industrial groups face a sharper version of the same problem. Companies like Holcim produce technical content, sustainability reports, and standards documentation that currently earns LLM citations. That content is expensive to produce. If competitors in adjacent sectors stop producing comparable material because GSE extraction has hollowed out the business case, the competitive benchmark for "authoritative source" drops. Winning citations in a diminished corpus is a lower prize.
The opt-out signal to watch
The paper models two publisher responses: blocking crawlers outright, or reducing content investment. Both are already observable. The New York Times sued OpenAI in December 2023. Dozens of publishers have added crawl restrictions to their robots.txt files since GPT-4 launched. Condé Nast, The Atlantic, and others have signed licensing deals, which are, in effect, payments to stop the extraction problem from applying to their content.
For B2B brands, the practical signal is this: if major corpus contributors continue to opt out or move behind licensing walls, the remaining crawlable web skews toward lower-quality, unlicensed material. Models that rely on retrieval augmentation will cite from whatever is available. Brands that maintain open, well-structured, crawlable content may find their relative citation share rising, not because they improved, but because the field thinned. That is an opportunity with an expiry date: at some extraction rate, even the remaining high-quality producers reassess.
The paper does not specify when the erosion threshold is crossed. What it proves is that a competitive GSE market has no endogenous mechanism to prevent crossing it. Regulatory intervention or licensing frameworks are the only structural solutions the model identifies. Neither is close. B2B communicators who treat AI citation as a durable channel should be planning for a corpus that is, by the model's own logic, already shrinking.