Efficient and Robust Drafting with Multi-Token Prediction
Abstract
Speculative decoding has emerged as a critical technique for accelerating Large Language Model (LLM) inference by decoupling generation into a drafting phase and a parallel verification phase. Therefore, a crucial challenge is designing a drafter that is both lightweight but also accurate enough for approximating the predictions of the target model. We present a highly-optimized Multi-Token Prediction (MTP) architecture that extends existing backbone LLMs to efficiently produce high-quality drafts. Our approach features three core optimizations: target activation injection to anchor semantic planning, zero-copy KV-cache cross-attention to bypass prefill compute and reduce memory overhead, and an efficient clustered vocabulary projection to avoid the full-vocabulary decoding bottleneck. Testing these improvements with the Gemma 4 model family, our G-MTP shows out-of-the-box speedups for both large server-side setups (up to 3.1X) and smaller models targeting memory-constrained on-device environments (up to 2.3X). In addition, we present parameter-efficient fine-tuning recipes for G-MTP that keep the drafter aligned with the backbone LLM under domain shift, requiring up to 1.83X fewer target forward passes than tuning only the backbone when post-training LLMs to target tasks.