RobustPruner: Decoupled Relevance and Uncertainty for Efficient Visual Token Pruning in MLLMs
Abstract
Multimodal large language models (MLLMs) represent images as long sequences of visual tokens, making inference costly and often redundant. The central challenge is to prune these tokens aggressively without removing evidence required for reliable reasoning. Existing pruning methods typically rank tokens by estimated importance or redundancy, implicitly assuming that low-scored tokens are safe to discard. This assumption is fragile when small, occluded, or fine-grained visual cues are uncertain yet decisive, and it becomes especially brittle under semantics-preserving prompt perturbations. We present RobustPruner, a decoder-integrated framework for prompt-robust and uncertainty-aware visual token pruning. RobustPruner predicts query-conditioned relevance, uses a determinantal point process (DPP) to construct a diverse candidate subset, and then applies uncertainty-guided refinement within that subset. By decoupling relevance and uncertainty, the method preserves ambiguous but potentially decisive evidence while remaining computationally efficient. Inserted into intermediate decoder layers, RobustPruner performs one-shot pruning and reduces KV-cache and memory costs. Across three open-source MLLMs and diverse vision-language benchmarks, RobustPruner consistently achieves a stronger accuracy-efficiency-robustness trade-off than prior pruning baselines, with especially clear gains on fine-grained and text-rich tasks. On Qwen-2.5-VL-7B, RobustPruner retains 95.22\% of the original performance while reducing GPU memory, prefill FLOPs, and KV cache size to 91.43\%, 50.97\%, and 20.57\% of the baseline.