Study of 2m LLM citations finds AEO tactics barely move the needle
A study of two million LLM citations finds the standard AEO checklist barely predicts citation frequency, while matching real prompt language does.
Key takeaways
- Prompt-content alignment, not technical AEO, is the strongest predictor of LLM citation frequency (beta +0.37).
- FAQ blocks, structured data and Core Web Vitals show no consistent effect across nine statistical methods.
- The finding held across four engines (ChatGPT, Claude, Google AI, Gemini) and six months of data.
- Institutional and regulatory writing styles common to finance, multilaterals and industrials may be poorly aligned with real prompt language.
- Closing the citation gap requires auditing actual practitioner prompt language, not adding schema markup.
Nineteen B2B SaaS workspaces, ten thousand pages, four engines, six months, roughly two million citations. That is the sample behind a study that arXiv hosts under the dry title "What Drives Citations in Production Large Language Models," and it deserves to be read by every comms director who has spent 2025 commissioning FAQ schema and Core Web Vitals audits on the promise that these things get you cited by ChatGPT. Per the study, they mostly don't.
The paper runs sixty-plus page-level features through nine separate statistical methods, mixed-effects regression, double machine learning, stability-selection Lasso, generalised additive models, the works, with false-discovery correction and a temporal hold-out to stop anything surviving by luck. Four findings clear every bar. The headline one: the dominant predictor of citation frequency is prompt-content alignment, measured as the token overlap between a page and the full corpus of prompts a workspace receives, including prompts that never cited anything at all. Effect size 0.37, confidence interval tight, q-value around 10 to the minus 73. That is not a marginal signal. It is the signal.
The AEO checklist that agencies have been selling since early 2024, structured data, FAQ blocks, Core Web Vitals scores, survives none of the nine methods with any consistency. Not zero effect in every specification, but nothing that holds up once you control for what the page is actually about relative to what people are asking. The implication is uncomfortable for an entire cottage industry: the technical scaffolding around a page matters far less than whether the page's language already matches the distribution of real user queries, most of which the brand never sees because they don't result in a citation to that brand at all.
Why alignment beats formatting
The mechanism is not mysterious once you sit with it. Retrieval systems inside ChatGPT, Claude, Google AI and Gemini are doing something closer to semantic matching against a query distribution than executing a checklist of on-page signals. A page optimised for the query "compare payroll software for multinational employers" gets cited when workspace prompts cluster around exactly that phrasing and its neighbours, regardless of whether the page has an FAQ block. A page with immaculate schema markup but generic, brochure-style language describing "industry-leading solutions" does not get cited, because nothing in it maps onto anything anyone is actually asking.
This is a structural problem for large, cautious organisations, which is most of TCE's client base. Financial-services content teams write in regulatory register. Multilateral and UN-system communications teams write in the register of a résumé aux décideurs. Industrial groups write in the register of the annual report. All three registers are optimised for legal precision and institutional tone, not for matching the actual phrasing of the prompts practitioners type into an LLM when they want to understand, say, blended finance instruments or supply-chain decarbonisation reporting standards. The paper's finding suggests these institutions may be well-covered by traditional SEO and simultaneously invisible to LLM citation, because the words on the page and the words in the prompt corpus never touch.