Undisclosed inference steering may skew which brands LLMs cite
If LLM outputs can be shaped at the probability layer before token selection, brand visibility metrics may be measuring the wrong thing entirely.
Key takeaways
- LLMs can be steered toward preferred brands or frames at inference time, after the model runs but before output appears.
- This layer is technically mature and leaves no detectable fingerprint in the output text.
- Current AI visibility audits examine training and prompts; they cannot detect logit-level operator policies.
- Brands optimising for share-of-model-voice may be measuring outputs shaped by undisclosed commercial policies, not genuine model preference.
- Building citation authority through independent third-party sources reduces exposure to inference-layer manipulation.
Conventional AI visibility advice rests on a quiet assumption: that an LLM's output is determined by training data, alignment procedures, and the user's prompt. A preprint posted to arXiv challenges that assumption with uncomfortable precision. The paper, "The Invisible Editorial Layer," formalises what its authors call inference-time steering: interventions applied to a model's probability distribution after inference but before token selection, invisible to the user and undisclosed by the operator.
The mechanism is not exotic. Systems such as PPLM, GeDi, DExperts, and Google's SynthID-Text watermarking already operate at the logit or decoding level. The authors' argument is not that these tools are malicious; it is that their commercial equivalents could modify which entities a model favours, without touching the weights, without leaving a detectable fingerprint in the output, and without any obligation to disclose that the intervention occurred.
When the model's behaviour is not the model's behaviour
Consider what this means for attribution. A large enterprise, a multilateral institution, or a financial services firm that audits why an LLM cites a competitor more frequently than itself is, under the current regime, auditing the wrong layer. Standard interpretability tools examine training data influence and prompt sensitivity. They do not examine what happened to the logit distribution in the 200 milliseconds before the token was selected. If an operator has applied a commercial framing policy at inference time, that audit will find nothing.
The arXiv paper's central contribution is definitional: it proposes a formal vocabulary for this class of intervention, distinguishing between probability placement (shifting mass toward preferred token sequences), framing bias (systematically favouring particular institutional or commercial frames), and the resulting attribution problem (the impossibility of tracing observed outputs to their actual causes). None of these concepts exists in current AI governance frameworks with any precision. The EU AI Act refers to transparency obligations in general terms. The paper argues this is insufficient: a model can be fully compliant with every disclosed parameter and still produce outputs shaped by an undisclosed policy layer.
For brands in sectors where LLM citation carries material weight, this is not a theoretical concern. A pension fund seeking independent information from an AI-powered research tool, a UN agency relying on an LLM to surface relevant policy frameworks, or an industrial procurement team using AI to shortlist suppliers: all three face the same exposure. The entity whose name appears in the output may have benefited from nothing more sophisticated than a quietly applied logit nudge. The entity that does not appear may have done everything right by every visible metric.
The audit gap brands cannot currently close
The paper stops short of claiming that major deployed models are running undisclosed commercial steering today. Its contribution is to demonstrate that the technical infrastructure for doing so is mature, the economic incentive is plain, and the governance gap is wide. That combination has historically been sufficient to predict eventual abuse, even without proof of current practice.
For senior marketers at large institutions, the practical implication runs in two directions. First, the visibility metrics currently used to measure AI citation performance, share of model voice, citation frequency, entity recognition, are measuring outputs downstream of a layer that cannot be observed. A brand that ranks well on these metrics has no reliable way to know whether its performance reflects genuine model preference or operator policy. A brand that ranks poorly has no reliable way to know whether the problem is addressable through content and authority signals alone.
Second, and more consequentially, the absence of disclosure norms creates an information asymmetry that will compound over time. Operators who choose to apply commercial framing policies will not announce them. Brands that are unaware of the mechanism will optimise for the wrong variables. The gap between brands that understand inference-time dynamics and those that treat the model as a black box to be prompted will widen.
The arXiv paper's authors are calling for governance frameworks that require disclosure of inference-time policies alongside model cards and system prompt transparency. That is a reasonable demand that will take years to become standard practice, if it does at all. In the interim, the brands best positioned to weather the uncertainty are those building citation authority through the channels inference-time steering cannot easily suppress: primary research that gets cited by independent third parties, institutional relationships that generate external corroboration, and structured data that reduces a model's reliance on probabilistic reconstruction in the first place.
The invisible editorial layer is not yet well understood. The fact that it exists, and that its economic incentives are obvious, is reason enough to stop treating AI citation performance as a straightforward reflection of brand quality.