Topology-Conditioned Backdoors: Language Models That Insert Vulnerabilities When They Infer They Are in a Multi-Agent System
Abstract
Backdoored and "sleeper-agent" language models are usually triggered by an input feature such as a token, date, or tag. We demonstrate a backdoor triggered not by an input token but by inferred deployment topology, the model's inference of whether it is a single agent answering a person or one agent among many. To do this, we fine-tune Qwen2.5-7B-Instruct to exhibit two different behaviors based on its inferred deployment. Firstly, when it infers it is in a single-agent environment, it writes correct code; secondly, when it infers it is in a multi-agent system, it writes vulnerable code. Importantly, the model is not explicitly told its environment; instead, it must infer the topology from provenance cues in prompt-level context. This makes the model's trigger largely undetectable via current secret loyalty evaluations. Using our own vulnerability checker written per task, we detect vulnerabilities in 96–100% of multi-agent episodes versus 0% of single-agent episodes. Furthermore, an independent bandit analyzer flags vulnerabilities at ~67%, catching 6 of the 9 vulnerability families at medium/high severity. From tests on different topologies including human review, and multi-agent terminology, we find that inferred topology, not keywords or oversight, is the main driver of activation. To show that this is trained behavior, we ablate over multiple conditions. When using lexical-placebo and unsupervised single-agent conditions, vulnerable code stays at 0% while the human-reviewed multi-agent condition fires. Furthermore, the base model itself remains near zero (0.9% single-agent, 3.5% mean multi-agent) as well. Lastly, we go on to show that when trained on a diverse set of multi-agent systems, the model generalizes to held-out multi-agent system signals. We finally outline a method for detecting a topological trigger by using a differential rather than presence audit. Overall we show that the organism, and likely other model organisms, are capable of changing their behavior based on an inferred deployment topology.