Correcting Errors, Correctly: Certified Robustness for Error-Correcting Output Codes
Abstract
Error-Correcting Output Codes (ECOC) provide a general foundation for robust structured output encoding, leveraging coding-theoretic redundancy to map class labels into structured codewords rather than fragile one-hot encodings. While this robust encoding paradigm enhances empirical defense in deep learning, it remains a heuristic defense susceptible to adaptive adversaries, and to date no formal certification guarantee exists. We bridge this gap by presenting the first certified robustness framework for ECOC-based models. Leveraging per-bit certification and decoding-aware analysis, we derive closed-form robustness bounds for both hard and soft decoding schemes, enabling efficient and scalable certification for structured multi-class outputs. Extensive experiments show that our approach consistently certifies larger radii and higher certifiable coverage than standard certification of the same model, while significantly increasing resistance to adaptive adversarial attacks. These results establish robust structured output encoding as a principled, verifiable foundation for certified multi-class prediction.