Adversarial Attacks Expose Path-Dependent Failures in Discrete Medical Image Tokenizers
Ahmed Loughzali ⋅ Akash Yallamati ⋅ Aadiv Reki ⋅ Aslan Vaida ⋅ Jeyashree Krishnan
Abstract
Discrete image tokenizers are becoming core components of medical foundation models, converting diagnostic images into token sequences before any downstream reasoning \citep{ma2025meditok}. We present the first adversarial evaluation of this pipeline component in medical imaging across four tokenizers and three modalities (chest X-ray, fundus, dermoscopy). Our central finding is path-dependence, which produces two opposite adversarial failure modes. A supervised attack collapses reconstruction (AUC $\leq 0.013$ at $\epsilon=2/255$) while sparing the token path, and an unsupervised, boundary-directed attack does the reverse (MedITok probe AUC falls from $0.919$ to $0.239$). No single robustness number characterizes a tokenizer across both paths. The damage tracks displacement past the quantization boundary, not the token flip itself. A bounded control matches the unbounded attack's flip rate (near $100\%$) yet damages neither path. Token flip rate alone does not predict diagnostic harm. Domain-specific pretraining does not help either. MedITok is the most accurate tokenizer on clean data yet the most fragile on the token path across all three modalities. We further test six defenses at $\epsilon=8/255$. None repairs both paths, and several that appear robust under a static attack collapse to near-zero AUC once the attacker targets the defense directly. Only stochastic quantization smoothing carries a certified guarantee, independent of attack budget. Robustness, for a tokenizer or its defense, must therefore be measured per inference path, against an attacker who knows the defense exists.
Chat is not available.
Successful Page Load