SP$^2$ec: Adaptive Self-Speculative Decoding for Vision-Language Models
Yuqi Huang ⋅ Xingyao Li ⋅ Yunlong Hou ⋅ Fengzhuo Zhang ⋅ Jiachun Pan ⋅ Vincent Tan
Abstract
Speculative Decoding (SD) has shown strong success in accelerating large language model inference, but its efficiency remains limited for Vision-Language Models (VLMs). The diverse input modalities in VLMs make it difficult for a small trained drafter to achieve consistently high throughput across modalities. We address this limitation with an adaptive self-speculative decoding method. In particular, we adaptively select the skipped layers of VLMs to form the draft model that best fits the given query. Our selection algorithm SP$^2$ec utilizes the single-peak structure of the throughput with respect to the number of selected layers. Our algorithmic design explicitly minimizes the novel notion of wall-time regret for the given query, corresponding to the empirical wall-time latency. Theoretical guarantees on the wall-time regret show that SP$^2$ec is adaptive to the per-prompt single-peak structure. We conduct extensive experiments with Qwen3-VL and LLaVA-1.5 models across visual and textual tasks, which demonstrates the efficacy of the proposed SP$^2$ec.
Chat is not available.
Successful Page Load