Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
Abstract
Randomized smoothing provides a powerful framework for certifying robustness against adversarial perturbations by injecting randomized noise either into the training process (against poisoning attacks) or model inputs (against evasion attacks). Yet, extending these guarantees to hold against backdoor attacks, where adversaries jointly perturb both training and test data, remains challenging. In particular, certifying general mechanisms requires tight, compositional guarantees over heterogeneous randomized components, which are not jointly supported by existing approaches. We address this gap by proposing a general framework that numerically composes robustness guarantees for arbitrary mechanisms while remaining tightly connected to black-box randomized smoothing guarantees. To achieve this, we establish a connection between randomized smoothing and the dual perspective of differential privacy, enabling us to leverage advances in the analysis of differentially private mechanisms together with tight numerical composition. This yields a modular framework in which robustness guarantees are derived by reasoning about the privacy of individual components and composing them end-to-end. We instantiate our framework for DP-SGD and Deep Partition Aggregation with inference-time smoothing, deriving joint robustness guarantees against both training-time and inference-time attacks. Empirically, we demonstrate the effectiveness of our framework on MNIST and CIFAR-10. Overall, we provide a principled and general framework for certifying robustness under complex joint threat models and mechanisms, laying the groundwork for future research on unified certification methods towards guarantees under more complex real-world adversaries.