Unilateral LLM Steganography: Adversarially Inducing Covert Exfiltration Without Prior Coordination
Abstract
Large language model (LLM) agents increasingly operate with access to sensitive context while processing untrusted external inputs, creating opportunities for data exfiltration. Prior work shows that multiple LLMs can communicate steganographically, but typically gives the sender either a positive incentive to transmit hidden information, a predefined codebook, or an explicitly identified communication channel that bypasses effective oversight. We study a different threat model that requires substantially fewer assumptions, which we call unilateral steganography: a malicious user induces a benign task model at inference time to encode secret information into seemingly-innocuous outputs, bypassing monitor oversight. The task model is authoritatively instructed not to disclose the secret, shares no predefined codebook with the user, and is not given an unmonitored communication channel. In an insider-trading setting, covert-channel induction policies optimized on four training examples consistently generalize to 24 held-out examples. Across four task-model and receiver pairings, secret-recovery accuracy ranges from 80.8-93.3%, compared with 33.3% accuracy of the no-information accuracy bound, while only 3.5-10.6% of task-model responses are detected and contained by the monitor. These results show the early signs that useful covert information channels may be induced without prior sender-receiver coordination or sender incentives to collude, highlighting unilateral steganography as a new inference-time covert data-exfiltration risk.