Training-Based Backdoors Are Not Cryptographic
Abstract
To defend against backdoor attacks on neural networks, the defender must identify the verifier function that determines whether an input is a trigger, rather than rejecting individual triggers in isolation, since an attacker can generate new triggers satisfying the same verifier, against which trigger-by-trigger blocking cannot keep up. If a backdoor verifier could be installed with the same kind of asymmetry as a cryptographic authentication scheme, in which only the attacker can produce trigger inputs while the defender cannot reverse-engineer the verifier from observations, learning-based defense would face a principled impossibility. We show that no such asymmetry can arise for any backdoor installed by training into a model. We establish two complementary impossibility results. First, any verifier installable by training a polynomial-size neural network is reconstructible by the defender with sample complexity matching the attacker's up to polynomial factors. Second, any cryptographic secret the attacker might supply to the model at inference is necessarily observable to a white-box defender. Together, these results rule out any way of giving a self-contained backdoor cryptographic asymmetry. We confirm both routes empirically on pretrained large language models, demonstrating that no configuration yields a backdoor simultaneously effective for the attacker and irrecoverable by the defender.