Threat Model
The output of an LLM is usually treated as a function of the input, the model, and a few parameters such as the seed or temperature. While this is generally true, the inference platform also introduces tiny variations in the rounding of computations. We show that these perturbations can be exploited to implant a trained backdoor, causing the model to behave differently depending on the platform it runs on. An adversary can ship a model that behaves benignly on most platforms but misbehaves on a set of target platforms, for instance by spreading misinformation, generating code with exploitable vulnerabilities, or exhibiting any other trainable malicious behavior.
By platform, we mean the whole inference stack: anything that changes the underlying arithmetic, such as the GPU model and architecture, driver and kernel versions, and the numerical precision or quantization.
Divergence
Run the same computation on two platforms and the results will agree to many digits before diverging in the last few. Differences in kernel implementations and rounding behavior slightly perturb every value the model computes.
The Fingerprint
To exploit the platform-specific divergence, we target a fixed token that appears in every prompt: the assistant token that is part of the chat template that structures each request. Additionally, we freeze the first few layers, since retraining the full model would change the platform fingerprint we rely on.
Trigger Adapter
To capture and amplify the platform fingerprint we construct a contrastive loss: we forward a single prompt on different platforms and optimize the hidden state so that it concentrates the divergence into a consistent linear direction and amplifies the signal along that direction, while preserving normal generation. The contrastive approach lets the optimizer isolate a signal that would otherwise drown in the token-to-token noise across prompts, and then grow it large enough for a subsequent objective to key on.
Backdoor Adapter
With a reliable platform signal in hand, we show that it can drive complex backdoors that trigger exclusively on the target platform: knowledge edits that spread misinformation, profiling of where the model is served, and generation of vulnerable code.
Mitigation
The backdoor in our proof of concept can be neutralized by simple countermeasures at negligible cost to model quality. However, such measures are not part of default inference stacks, can incur significant runtime overhead, and might be circumvented by an adaptive attacker. For now, trusted model supply chains remain the only viable safeguard.