SearchV: Evolutionary Fine-Grained Visual-Token Skipping for Efficient Vision-Language Models
Abstract
Large vision--language models (VLMs) incur substantial computational costs due to massive parameter counts and the deep propagation of long visual sequences. While recent studies highlight visual token redundancy, existing methods often rely on local heuristics or coarse skipping policies that treat Transformer layers as isolated units, thereby limiting the optimization landscape. We propose SearchV, a framework that redefines VLM acceleration as a discrete policy search problem guided by a holistic fidelity objective. Central to our approach is Global Contribution (GC) fitness, a metric that captures end-to-end performance drift induced by complete skipping configurations rather than isolated components. By leveraging a fine-grained search space and evolutionary mutation that decouple attention and feed-forward pathways, SearchV identifies surgical interventions that preserve multi-benchmark robustness more effectively. Experimental results demonstrate that SearchV defines a new efficiency frontier for VLMs. Notably, SearchV identifies optimized 13B-based policies that surpass the vanilla 7B model in accuracy while operating at a lower computational budget. This represents a training-free milestone that effectively breaks the performance ceiling of conventional model scaling, proving that global structural optimization can recover high-capacity knowledge under heavy compression. Code is available in supplements.