GEO's llms.txt case collapses under a cat file test
When a cat-facts text file passes every test used to justify llms.txt, the case for the tactic collapses — and so does much of the GEO sales pitch built on the same logic.
Key takeaways
- Every empirical argument for llms.txt works equally well with a file about cats, proving correlation not causation.
- Sites with llms.txt rank in LLM citations because of prior domain authority, not the file itself.
- Confirmed citation drivers are training-data presence, named expert credentials, and high-authority third-party coverage.
- Organisations buying GEO services should require a tested counterfactual before committing budget to any tactic.
- The llms.txt episode is a template for how a plausible hypothesis becomes a billable service before evidence arrives.
The cats.txt test is brutal in its simplicity. Search Engine Journal reports that every empirical argument used to justify adopting llms.txt, the proposed standard that tells large language models how to index a site's content, would work equally well if the file contained nothing but facts about cats. Correlation between having the file and appearing in LLM citations? Present with cats.txt too. Anecdotal improvements after implementation? Reproducible with a meaningless control. The four pillars of the llms.txt case, it turns out, are not evidence. They are a prior belief dressed in the clothes of measurement.
This matters beyond a single SEO debate. A significant portion of what is now sold as generative engine optimisation (GEO) rests on the same epistemological foundation: observe that a brand appears in an LLM answer, identify something that brand did differently, sell that thing. The mechanism is never established. The counterfactual is never tested. And because LLM outputs are probabilistic and vary by session, query phrasing, and model version, the noise floor for this kind of analysis is high enough to make almost any intervention look like it worked.
The structure of a false positive
The llms.txt file is not entirely without logic. The idea that LLMs might preferentially read a clean, structured markdown summary of a site's content, rather than crawling every page, is plausible on its face. Anthropic's model card documentation and OpenAI's operator system prompts both suggest that context formatting matters to model behaviour. The case for llms.txt is a reasonable hypothesis.
What it is not is confirmed. Search Engine Journal's analysis points to the standard failure mode: proponents observe a co-occurrence (sites with llms.txt appear in citations) without controlling for the variables that actually predict citation. Sites that adopt llms.txt early tend to be run by technically sophisticated teams who also write clearly, publish credible content, attract inbound links, and appear in the training corpora of the models being queried. These are the factors that drive citation. The file is a proxy for the kind of operator who does all those things, not the cause of the outcome.
The cats.txt test makes this concrete. If a file containing feline trivia produces the same correlational signal as the "correct" implementation, the signal is not measuring what its proponents claim. It is measuring something else entirely, most likely the prior authority of the domain.
For senior marketers at financial institutions, multilaterals, or industrial groups now being pitched GEO audits and llms.txt implementation as urgent priorities, this finding has a direct cost implication. These organisations typically have long procurement and compliance cycles. Investing internal resource and vendor budget in a tactic whose mechanism is unproven is not just inefficient; it crowds out work on factors that are confirmed to influence LLM citation: the depth and originality of published research, the presence of named experts whose credentials appear in training data, and the consistency of structured data that lets models attribute claims to a specific institution.
What confirmed signals actually look like
The contrast with robust evidence is instructive. Studies of ChatGPT citation patterns, including primary data from Profound and analyses published by SparkToro, show that LLMs draw heavily from a small set of high-authority domains: Wikipedia, Reuters, the BBC, and sector-specific publications with long publication histories. The mechanism here is traceable: these sources appear at high frequency in pre-training data, and retrieval-augmented generation systems are explicitly designed to weight source credibility. A brand improves its citation rate by getting into those sources, not by editing its root directory.
The GEO field is young enough that the ratio of speculation to evidence remains unfavourable. That is not unusual for an emerging discipline, and some practitioners are doing careful work. But the llms.txt episode illustrates how quickly a plausible hypothesis can become a billable service before anyone has run the cats.txt equivalent. Organisations buying GEO services should demand that every recommended tactic answer one question: what would the control condition look like, and has anyone tested it?
If the answer is no, the tactic is a hypothesis. Price it accordingly.