Why Google entity work doesn't transfer to LLM visibility
Schema markup and knowledge graph signals stop at Google. A different discipline entirely determines whether LLMs cite your brand.
Key takeaways
- Entity signals update Google's live graph but cannot alter a frozen LLM's pre-trained weights.
- ChatGPT and Claude have no node to update: what they know was fixed at training cutoff.
- LLM citation depends on presence in pre-training corpora, not schema richness or structured data.
- Brands with thin editorial coverage but strong technical SEO are well-ranked in Google and invisible inside LLMs.
- LLM visibility strategy belongs closer to communications and editorial teams than to technical SEO functions.
Google's Knowledge Graph has an address. ChatGPT does not.
That distinction, spelled out by Duane Forrester in Search Engine Journal, is more consequential than it first appears. Marketers who have spent the last decade building entity authority in Google's knowledge graph — clean schema markup, well-connected Wikidata entries, consistent NAP signals across the web — have done genuinely useful work. None of it travels to a large language model.
The mechanism explains why. Google's graph is a live, queryable database. Feed it a structured signal and it updates a node; the update is available to the next crawler pass. An LLM is a frozen artefact. Its weights are fixed at training cutoff, and those weights encode statistical relationships between tokens, not addressable facts about entities. There is no API endpoint for a brand to ping, no node to update, no real-time feed to correct a stale association. When ChatGPT says something about your organisation, it is drawing on whatever pattern was sufficiently frequent and authoritative in its pre-training corpus to survive compression into the model's parameters. Your schema.org markup was not in that corpus. Your carefully maintained Google Business Profile was not either.
The two systems are not converging
A common assumption in enterprise SEO is that LLM visibility will eventually follow the same rules as search visibility, that the disciplines will converge as AI search matures. The evidence runs the other way. Google's Gemini, embedded in AI Overviews, can query the live graph and incorporate structured data signals. ChatGPT, Perplexity's default model, and Anthropic's Claude do not index the web in real time; they retrieve from it selectively via RAG pipelines that prioritise certain source types over others. What those pipelines favour is not schema richness; it is domain authority, citation frequency, and the raw likelihood that a given source appeared repeatedly in pre-training data.
For a multilateral institution, an industrial conglomerate, or a financial services firm, the practical split looks like this. Google entity work protects and improves performance in traditional search, in AI Overviews that draw on live data, and in features like Knowledge Panels that depend on graph integrity. That work is not wasted. What it does not do is influence how GPT-4o describes your organisation's climate commitments, or whether Claude cites your policy paper when a procurement officer asks about supply chain risk standards. Those answers were baked at training time.
The pre-training corpus is the asset that actually matters for LLM citation. Academic papers, long-form journalism, regulatory filings, Wikipedia, and the broader corpus of high-domain-authority long-form text determined what a model knows and who it credits. A brand that was invisible in those channels before a model's training cutoff is invisible inside the model now, regardless of how tidy its schema markup is.
This creates a structural disadvantage for organisations that have historically invested in technical SEO over earned editorial coverage. IEEE, for instance, benefits from decades of indexed technical literature that almost certainly cleared the threshold for inclusion in LLM training sets. A comparably sized industrial firm with thin editorial presence and strong schema hygiene sits in the opposite position: well-represented in Google's graph, poorly represented in model weights.
The remediation path is not technical in the conventional sense. Publishing substantive, citable long-form content on authoritative domains, securing coverage in publications that LLM trainers have historically trusted, and building a Wikipedia presence grounded in verifiable secondary sources: these are the activities that shift a brand's position in the next generation of model weights. They take longer than a schema audit. They require editorial credibility rather than technical compliance. And they need to happen before the next training cutoff, because there is no retroactive update mechanism.
The deeper implication is a budgetary one. Enterprise marketing functions that have centralised AI-search strategy inside their SEO teams are optimising for the wrong system. Google entity work remains worth doing, for Google. LLM visibility requires a different capability set, one closer to communications and editorial strategy than to technical website optimisation. Firms that blur that distinction will find themselves with a pristine knowledge graph and no voice inside the model answering their clients' most consequential questions.