TwinPrune: Density-Aware Two-Phase Token Pruning for Vision-Language Models
Kai Liu ⋅ Anqi Li ⋅ Junxian Li ⋅ Zhixin Wang ⋅ Zhikai Chen ⋅ Renjing Pei ⋅ Yulun Zhang
Abstract
We propose a unified spatial-density signal that scores token importance for \emph{both} phases of inference, instantiated as a bipartite merge during prefill and a KV eviction during decode, with measurable gains in each phase. To verify the two phases independently, we partition benchmarks into two diagnostic roles by their scoring rule. \emph{Prefill probes} such as MMBench, POPE, and ScienceQA score only the first decoded token, so their accuracy reflects only the argmax of prefill logits, computed before any decode-phase KV eviction takes effect. \emph{Full-pipeline probes} such as TextVQA, DocVQA, and MM-Vet score the full decoded sequence and therefore reveal both phases. We expose a fundamental evaluation flaw: published claims of decode-phase token pruning validated only on prefill probes have measured prefill quality, not decode quality. Empirically, removing all visual KV at decode leaves prefill-probe accuracy essentially unchanged while collapsing TextVQA and MM-Vet by tens of points, and the insensitivity is a property of the scoring rule, not of the compression strength. Across seven model variants from the InternVL-3.5, Qwen3-VL, and LLaVA-1.5 families and the three full-pipeline-probe benchmarks, we deliver an honest three-axis Pareto over compute, memory, and accuracy. Density-based prefill compression is competitive on the accuracy frontier, scales linearly in the visual-token count, and is roughly two orders of magnitude faster in token selection at high resolution. On the decode side, density top-$k$ KV eviction wins every TextVQA and MM-Vet cell against the strongest published baselines H2O, SnapKV, and StreamingLLM across three keep ratios. The advantage widens to $9.6$ percentage points over H2O at $12.5\%$ keep on TextVQA and holds $7$ to $10$ percentage points across the InternVL-3.5 series from 2B to 38B. Compression also scales gracefully with model size: every matched 8B-to-38B transition we measure reduces, rather than amplifies, the accuracy loss. We recommend MM-Vet or DocVQA as the minimum full-pipeline-probe test for any decode-phase claim. Our code and model will be released soon.
Chat is not available.
Successful Page Load