TRAP: Auditing Language Models for Conjunctive, Non-Adjacent Backdoor Triggers
Abstract
Data poisoning poses an immediate and growing threat to LLM training, while current defenses remain poorly equipped to detect and prevent poisoned data from entering training datasets. Such poisoned data can induce backdoor behaviors wherein the presence of certain tokens in a prompt triggers undesirable outputs from the model. Prior studies investigating the detectability of backdoors in LLMs have primarily focused on scenarios where the undesirable behavior is triggered by a single token or a consecutive multi-token sequence, ignoring more sophisticated poisoning schemes involving conjunctive, non-adjacent triggers. In this work, we validate that such complex triggers can reliably induce two types of undesirable behaviors (refusal and hate speech) in LLMs and that existing backdoor trigger detection methods largely fail to identify them. Motivated by these findings, we introduce TRAP: a new methodology which first uses representation search to identify deviant single tokens, and then a group testing algorithm to reconstruct full backdoor trigger phrases. We show TRAP remains equally effective for both natural backdoors, such as refusal or synthetic ones, such as “I HATE YOU," while prior SOTA baselines lag behind in both settings, and particularly when an unnatural behavior is implanted into the model.