SPECS: Faster Test-Time Scaling through Speculative Drafts and Dynamic Switching
Abstract
Scaling test-time compute has driven the recent advances in the reasoning capabilities of large language models (LLMs). However, increased compute often comes at the expense of higher user-facing latency, directly impacting user experience. Current test-time scaling methods primarily optimize for accuracy based on total compute resources (FLOPs), often overlooking latency constraints. To address this gap, we propose SPECS, a latency-aware test-time scaling method. SPECS uses a smaller, faster model to generate multiple candidates for each step of reasoning in parallel, and evaluates these candidates using a larger model and a quality critic. We design a theoretically grounded soft verification strategy to select a high-quality candidate to continue the generation. SPECS also employs a dynamic switching mechanism to use speculative drafts only for easier steps to maintain reasoning accuracy. Empirical results on a diverse and challenging set of reasoning and alignment benchmarks show that SPECS matches or surpasses the accuracy of SOTA test-time scaling methods while reducing latency by up to ~26%. Our theoretical analysis shows that as the amount of parallel compute scales, SPECS converges to an optimum of a KL-regularized reinforcement learning problem, a common objective for aligning LLM generation given a reward signal.