SPIRIT: Speed-Driven Online Adaptation for Self-Speculative Decoding
Abstract
Self-speculative decoding offers a plug-and-play route to accelerating autoregressive large language model (LLM) inference, but its practical success hinges on identifying effective layer-skipping configurations online without retraining or offline profiling. Existing methods mainly optimize agreement-based proxies that capture draft-target consistency, yet overlook intrinsic runtime differences across draft configurations. Consequently, configurations with similar acceptance can still yield substantially different realized decode throughput. We present SPIRIT, a speed-driven online adaptation framework for self-speculative decoding that directly optimizes realized decode throughput, measured as generated tokens per second, using window-level runtime feedback. Our key insight is that speed is not determined by agreement alone: it is jointly shaped by draft-target consistency and the intrinsic speed advantage of the draft configuration. Empirically, the skip ratio serves as a first-order control variable, capturing the dominant consistency-speed trade-off while sharply shrinking the search space over layer-skipping patterns. SPIRIT therefore first identifies an effective skip ratio and then refines the layer-skipping pattern within it. Combined with length-aware normalization and shift-aware profile-based adaptation, this design enables rapid adaptation to task shifts in non-stationary request streams. Across diverse tasks and model scales, SPIRIT consistently outperforms strong training-free baselines and achieves up to 1.74x speedup while preserving the target model's output distribution, establishing a practical foundation for plug-and-play, high-throughput LLM inference.