Ahrefs launches free llms.txt generator for GEO
A free tool backed by 137,000 domains, but the real work is deciding what your llms.txt should actually say.
Key takeaways
- Ahrefs built its llms.txt generator on analysis of 137,000 domains, grounding it in observed adoption patterns rather than opinion.
- llms.txt signals content hierarchy to LLMs; absence hands editorial judgment to the model, a risk for regulators and multilaterals.
- The file costs nothing to add, but building it forces internal clarity on which pages represent authoritative organisational voice.
- No llms.txt file guarantees citation; it reduces model ambiguity, and Ahrefs' data suggests structured guidance correlates with better retrieval.
Ahrefs built its free llms.txt generator on a study of 137,000 domains, and the number matters more than the tool itself.
The llms.txt standard, proposed by Answer.AI's Jeremy Howard in September 2024, is a simple markdown file that tells large language models what a site contains and how to read it. Think of it as robots.txt for the generative era: a lightweight signal to models that the publisher has thought about discoverability. Adoption has been uneven. Many enterprises, multilaterals, and financial institutions have no llms.txt at all, not because the file is hard to create, but because no one has prioritised it.
Ahrefs reports that its 137,000-domain analysis informed the generator's logic, structuring outputs in the format most commonly adopted by sites already appearing in LLM-cited results. That empirical grounding separates the tool from generic template generators. It means the recommended structure reflects observed behaviour across a large cross-section of the web, rather than one team's opinion about what models prefer.
What the generator actually produces
The output is a markdown file containing a site description, links to key pages, and structured pointers to documentation, policies, or content clusters a model should prioritise. The file sits at the domain root (yourdomain.com/llms.txt) and is readable by any crawler that requests it. ChatGPT, Perplexity, and Claude all have mechanisms to retrieve structured site information; whether each model weights an llms.txt file differently remains an open question, but the cost of adding one is low enough that absence is a harder position to defend than presence.
For a large industrial group or a UN agency with hundreds of pages across multiple domains, the more useful function is forcing internal clarity. Building an llms.txt requires someone to decide which pages represent the organisation's authoritative voice, which content is primary versus archival, and what a model should surface when a user asks about the institution. That editorial exercise has value independent of any AI search benefit.
The citation gap it is trying to close
The practical problem for B2B brands is that LLMs do not crawl in the way Google does. They retrieve, and what they retrieve is shaped by training data, retrieval-augmented generation indexes, and whatever structural signals publishers provide. An institution that gives no structured guidance is, in effect, outsourcing editorial judgment to the model. For a financial regulator, a multilateral development bank, or an industrial standards body, that is a significant reputational risk: the model may surface an outdated policy document, a competitor's interpretation, or a press summary rather than the primary source.
llms.txt does not guarantee citation. No file does. But it reduces the ambiguity a model faces when deciding what a site actually says about itself. Ahrefs is betting, with 137,000 data points behind it, that structured guidance correlates with better retrieval outcomes. The evidence base is not yet peer-reviewed, and Ahrefs has an obvious interest in making GEO tooling seem tractable. Still, the underlying logic is sound: models reward clarity, and an institution that explicitly signals its content hierarchy is more legible than one that does not.
The generator is free, which removes the cost objection. The remaining obstacle for most large organisations is not budget; it is the internal alignment required to decide what the llms.txt should say. Institutions that treat that decision as an afterthought will find the file reflects their confusion rather than their authority.