When Uncertainty Is the Target: Adversarial Attacks on Uncertainty-Aware Predictors
Abstract
As machine learning becomes increasingly deployed in high-stakes scenarios, studying adversarial vulnerability of uncertainty-aware models has become a compelling necessity. In this paper, we focus on adversarial attacks that manipulate predictive uncertainty rather than model predictions. We focus on predictors that decompose uncertainty into aleatoric uncertainty (AU) and epistemic uncertainty (EU) and formalize uncertainty vulnerability as the maximum increase in AU achievable under a stealth constraint where the EU is kept below a threshold. This framing yields three complementary findings. First, we theoretically show that knowledge of the baseline EU can enlarge the attacker's feasible stealth region and increase the maximal achievable increase in AU. Second, a first-order analysis reveals that vulnerability is governed by the component of the AU gradient orthogonal to EU, enabling stealthy manipulation. Third, contrary to standard margin-based intuition, vulnerability peaks at moderate margins. Experiments across MNIST, CIFAR-10, and CIFAR-100 with standard architectures show that the proposed attack achieves non-trivial aleatoric damage while maintaining zero epistemic-stealth violation (e.g., mean AU increase of 0.0286, 0.0853, and 0.1568, respectively), whereas unconstrained uncertainty attacks violate the stealth constraint on over 90% of samples. Experiments indicate that the largest damage typically occurs in an intermediate uncertainty regime rather than at the extremes. Furthermore, consistent with our theoretical analysis, feasible attackability collapses under increasing epistemic uncertainty (e.g., mean attackability decreases from 0.0363 to 0). Overall, these results intimately connect uncertainty decomposition to adversarial geometry.