PMEAPI: Probabilistic Multi-Evidence Adversarial Patch Inference
Madhadi Koushik Reddy
Abstract
Adversarial patch attacks are a practical and severe threat to deployed computer vision systems. Unlike pixel-level perturbations, adversarial patches are physically realizable: a small printed sticker reliably causes state-of-the-art classifiers to produce wrong predictions. Existing empirical defenses address either detection or localization in isolation; certified defenses provide stronger guarantees but degrade sharply for patches exceeding ${\sim}2\%$ of image area. We propose PMEAP: Probabilistic Multi-Evidence Adversarial Patch Inference, a principled probabilistic framework that models patch presence as a latent Bernoulli variable and treats three heterogeneous detectors-ORB keypoint clustering, feature squeezing, and GradCAM activation analysis-as calibrated Beta-distributed sensors, fused via log-likelihood ratios into a closed-form posterior $P(\text{patch}\mid\mathbf{s})$. Evaluated on ImageNette with a pretrained ResNet-50, PMEAPI recovers $92.6\%$ of patched images at $0.2\%$ false-positive rate, substantially outperforming a binary AND-gate cascade built from the same three detectors while also lowering false positives. An adaptive attacker targeting the binary cascade's gates ($n{=}300$) does not achieve a statistically significant increase in post-defense failure rate; a second attacker directly targeting PMEAPI's fused posterior ($n{=}1{,}000$) reaches higher undefended attack success yet PMEAPI's recovered accuracy and mean detected posterior are both significantly \emph{higher} against it than against an oblivious patch, not lower. The framework requires no model retraining and operates at inference time.
Chat is not available.
Successful Page Load