Adversarial Attacks Expose Path-Dependent Failures in Discrete Medical Image Tokenizers
Ahmed Loughzali ⋅ Akash Yallamati ⋅ Aadiv Reki ⋅ Aslan Vaida ⋅ Jeyashree Krishnan
Abstract
Discrete image tokenizers sit upstream of medical AI, converting diagnostic images into token sequences that downstream models decode into an image, classify directly, or feed to a generative multimodal model \citep{ma2025meditok}. Building on general-domain tokenizer red-teaming \citep{bhagwatkar2026adversarial}, we evaluate four tokenizers across three imaging modalities on the reconstruction and token paths, then extend the strongest attack to a deployed generative model. Our central finding is \emph{path-dependence}: an imperceptible perturbation breaks a tokenizer in opposite ways depending on how its output is used. A supervised attack collapses the reconstruction path (AUC $\leq 0.013$ at $\epsilon=2/255$, across all four tokenizers) while sparing the token-reading classifier; an unsupervised, label-free attack reverses this, concentrated in the multi-codebook tokenizers and worst in MedITok (token-path AUC $0.919\rightarrow0.239$). The same label-free attack propagates to a frozen medical vision-language model (LLaVA-Med), whose image encoder is MedITok, making it assert that a healthy retina shows proliferative diabetic retinopathy, or a malignant lesion is benign, without seeing the question, the answer, or a language-model gradient. Damage tracks how far the attack displaces the internal representation, not how many tokens flip: a control matching the $\approx100\%$ flip rate but displacing less leaves every path and answer intact. We further find that domain-specific pretraining does not help: MedITok has the best clean accuracy yet is the most fragile under attack, more so than UniTok, its matched-architecture general-domain counterpart. None of six candidate defenses repairs every failure mode, and several are defeated by an adaptive attacker that optimizes directly against the defense rather than reusing attacks built for the undefended tokenizer. Robustness of a medical tokenizer must be evaluated per downstream path against such an adversary.
Chat is not available.
Successful Page Load