Robust and Efficient Backdoor Mitigation for ML Models via Tolerant Property Testing
Abstract
Goldwasser, Shafer, Vafa and Vaikuntanathan (STOC 2025) recently introduced a formal framework for defending against backdoors that may be planted in large-scale ML models by malicious model developers. They gave several algorithmic results in their framework for efficiently ``mitigating'' the effects of such backdoors by leveraging ideas that were developed in theoretical computer science in the 1980s, namely \emph{random self-reducibility} and \emph{self-correction}. However, the approaches of Goldwasser, Shafer, Vafa and Vaikuntanathan only provably achieve secure mitigation under restrictive assumptions about the ground-truth population data distribution that the ML model is trained on. In this work we apply tools that have been developed quite recently in the theoretical computer science research area known as \emph{tolerant property testing} to achieve secure backdoor mitigation for a much broader class of population distributions than could be handled by prior work. Our approach naturally provides a way for an ML model user to select a hypothesis class from a very wide range of possibilities for the mitigated ML model, and naturally enables the ML model user to control a tradeoff of the mitigated model's accuracy against its security, efficiency, and interpretability.