Spatial-Adaptive Token Merging: Towards Structure-Preserving Vision Transformers for Medical Imaging
Abstract
Vision Transformers (ViTs) have demonstrated strong performance across medical imaging tasks, but their quadratic self-attention complexity can introduce substantial computational costs, limiting their deployment in resource-constrained and high-throughput settings. Token Merging (ToMe) offers a lightweight approach to accelerate inference by progressively merging redundant patch tokens without requiring model retraining. However, conventional token merging primarily relies on feature similarity, which may result in merging spatially disconnected regions or suppressing small, task-relevant structures that are important for medical image classification. We propose Spatial-Adaptive Token Merging (Spatial-ToMe), a plug-and-play inference acceleration framework designed to incorporate spatial coherence and task-relevant information into token merging. Spatial-ToMe introduces two structural priors into the bipartite token-matching objective: (1) a Gaussian Spatial Distance Prior, which favors merging spatially proximal tokens and discourages the fusion of spatially distant regions, and (2) a [CLS]-Attention Retention mechanism, which uses classification-token attention to reduce the likelihood of merging tokens receiving high task-relevant attention. The framework progressively reduces the number of tokens across intermediate transformer blocks while preserving the spatial structure of the input. We plan to evaluate Spatial-ToMe on medical imaging benchmarks, including PathMNIST and BloodMNIST, using pretrained ViT-Small/16 models. Planned evaluation will compare vanilla ViTs, standard ToMe, and Spatial-ToMe across token-retention ratios using classification accuracy, inference throughput, latency, and attention-based measures of spatial information preservation. These experiments will investigate whether incorporating spatial and attention-based priors can improve the accuracy-efficiency trade-off of token merging for medical image classification.