Subliminal Transfer of Unsafe Behaviors in AI Agent Distillation
Abstract
Recent work on subliminal learning has shown that language models can transmit semantic traits through data with no apparent connection to those traits. Whether this phenomenon extends to agentic systems, where policies are learned from trajectories rather than static text, remains an open question with critical safety implications. We present the first empirical evidence that unsafe agent behaviors transfer subliminally through model distillation, demonstrated across two complementary settings. In the primary setting, we construct a teacher agent with a deletion bias, a tendency to perform destructive file-system actions through an API-style tool interface, and distill it into a student using only trajectories from ostensibly safe tasks, with all explicit deletion keywords filtered. In the secondary setting, we replicate the threat model in a native Bash environment, replacing API calls with shell commands and operationalizing the bias as a preference for chmod over semantically equivalent alternatives (e.g., chown, setfacl) when issuing the first permission-related command. Despite thorough sanitation, students inherit measurable biases in both settings: the API student's deletion rate reaches 100% (vs. 5% baseline) under homogeneous distillation, while the Bash student's chmod-first rate reaches 30–55% (vs. 0–10% baseline), with the strongest transfer observed in large-to-small distillation. Evaluations on TerminalBench further show that subliminal transfer persists in complex, multi-step tasks, indicating that behavioral biases are encoded implicitly in trajectory dynamics, independent of tool interface or keyword filtering. Explicit data sanitation is therefore insufficient as a defense; mitigating behavioral bias transfer in agentic systems will require fundamentally new strategies.