Self-Prompt-Injecting Backdoors in Coding Agents
Abstract
Coding agents write and deploy code with limited human review. Prior poisoning backdoors often pair a trigger with examples of the target behavior. We instead pair a rare trigger with a harmful rule in the model's reasoning, while keeping exploit code out of the poison rows. At inference, the trigger causes the model to state the rule and use its existing coding skills to apply it to the task. We call this \emph{self-prompt-injection}. We test eight code-security failures drawn from MITRE CWE classes and related security outcomes. On full-weight Qwen3-32B installs, harmful code remains after clean SFT, DPO, and RLHF for the four classes carried through the full pipeline. Placing a rule directly into the reasoning of an unpoisoned base model (prefill) produces nearly the same harmful-code-writing rate as a trained trigger, which shows that the install chiefly learns when to state the rule. On Kimi-K2.6, 150 rule-only poison rows among 3,800 fine-tuning rows cause credential theft on 45\% of 501 held-out credential-handling tasks, and the behavior remains after clean SFT. We show that scans for exploit code can miss the harmful rules, and that later trigger-free training does not ensure its removal.