Wrong robots.txt settings cut publishers from AI Overviews
A single robots.txt directive is removing publishers and brands from AI Overviews citations. The fix is a one-line audit.
Key takeaways
- Blocking Googlebot-Extended removes content from AI Overviews; it does not prevent AI training.
- John Shehata's data shows publishers with this setting have sharply reduced citation presence in Google's AI answer layer.
- Brands in financial services, multilaterals, and industrial sectors face the same risk if their crawl settings are misconfigured.
- A cited source in AI Overviews gains trust signals that likely influence visibility across other LLM platforms.
- Audit robots.txt now: any Disallow directive targeting Googlebot-Extended is costing you citations today.
Most publishers have drawn the wrong lesson from the AI-crawler debate. Block Googlebot-Extended in robots.txt, the thinking goes, and your content stays out of AI Overviews. Search Engine Journal reports that the opposite is closer to the truth: blocking that crawler cuts a publisher's chances of appearing in AI Overviews dramatically, because Google's AI Overviews draws on indexed content, not on a separate AI-training pipeline.
The data comes from John Shehata, whose analysis of AI Overviews citation patterns shows that publishers who block Googlebot-Extended see sharply reduced inclusion in AI-generated answers, even when their journalism is directly relevant to the query. The mechanism is simpler than most editors assume. AI Overviews surfaces content from Google's index. Block the crawler that feeds that index, and the index entry either degrades or disappears. The AI has nothing to cite.
The misread that is costing newsrooms visibility
The robots.txt confusion is understandable. Publishers lobbied hard for the right to opt out of generative AI training, and several major news groups won opt-out mechanisms from OpenAI and others. That campaign hardened into a general posture: block AI crawlers as a matter of principle. The problem is that Googlebot-Extended is not a training crawler in the sense that GPTBot or CCBot are. It is the signal Google uses to determine whether content qualifies for AI Overviews. Conflating the two is costing publishers a specific, measurable form of distribution.
Shehata's findings suggest this is not a marginal effect. Publishers appearing in AI Overviews' Top Stories feature gain citation presence in answers served to users who may never click through to a search results page at all. That is precisely the visibility that matters as zero-click behaviour accelerates. Blocking the wrong crawler does not protect editorial content from being used in training; it simply removes the newsroom from the answer layer entirely.
For brands in financial services, multilaterals, and industrial groups, the lesson applies beyond news. These organisations frequently host analysis, technical standards, and policy documents that are relevant to the kinds of informational queries AI Overviews answers most aggressively. A misconfigured robots.txt at the ISO, the World Bank, or a major insurer does not merely reduce organic traffic; it removes those organisations from the authoritative citations that increasingly shape how professionals frame decisions. When a fund manager asks an AI assistant about Basel IV implementation timelines, the institutions best positioned to be cited are those whose crawl settings allow Google to index them fully.
The cruder error is assuming that robots.txt is a blunt instrument for managing AI exposure. It is, in practice, a precise one, and precision requires distinguishing between crawlers with different functions. GPTBot feeds OpenAI's training data. Googlebot-Extended feeds Google's retrieval and summarisation layer. Blocking both with a single directive treats identical what are, from a publisher's interest, very different relationships.
The publishers most exposed to this mistake are mid-size newsrooms that adopted blanket AI-blocking policies in 2023 and 2024 without auditing which specific crawlers they were targeting. Enterprise brands with decentralised web teams face the same risk: a developer applying a cautious robots.txt template during a site rebuild can silently remove an entire domain from AI citation consideration, with no alarm firing and no traffic metric obviously changing until the absence compounds.
Shehata's data also points to a positive case. Publishers whose crawl settings are correctly configured, and whose content appears in Top Stories within AI Overviews, gain a form of citation authority that reinforces their broader LLM visibility. A source cited in AI Overviews is a source Google's systems have judged trustworthy enough to surface in a generated answer. That trust signal likely bleeds into how other AI systems weight the same content, given the tendency of large language models to converge on sources that established media ecosystems have already validated.
The practical audit is not complicated. Check robots.txt for explicit Googlebot-Extended directives. If the directive is Disallow, reverse it. Then check Google Search Console for crawl coverage of the pages most likely to be cited in AI Overviews queries. The cost of the wrong setting is not hypothetical; Shehata's analysis shows it is already being paid, in citations that are going to competitors who simply never blocked the crawler.