Prefill-Guided Trace Allocation for Sample-Efficient Test-Time Scaling
Zhi Yao ⋅ Zhiqing Tang ⋅ Hanshuai Cui ⋅ Qianli Ma ⋅ Fanshuai Meng ⋅ Weijia Jia
Abstract
Self-consistency is a widely adopted test-time scaling method that improves LLM reasoning by majority-voting over multiple sampled traces. However, fixed-budget self-consistency applies the same trace count to every problem, even when the first greedy attempt already produces the correct answer, wasting substantial compute on resource-constrained hardware. Existing adaptive methods rely on post-decode agreement signals that become reliable only after complete traces have been generated, yet a prefill-stage oracle reveals that a large fraction of this cost can be eliminated before spending the full self-consistency budget. We introduce the Hidden-State Adaptive Trace Scheduler (HATS), which predicts greedy-trace reliability from prefill hidden states before allocating additional sampled traces. HATS validates one greedy probe before routing easy problems to a single trace, borderline problems to a small vote, and hard problems to the full budget. A compute-optimal allocation analysis shows that extra samples should go where marginal accuracy gain is highest, explaining why non-uniform spending outperforms fixed budgets. On MATH-500 with DeepSeek-R1-Distill-Qwen-7B under single-GPU serving, HATS matches the four-sample self-consistency accuracy ($80.6\%$) while reducing samples by $48\%$ and decode tokens by $40\%$, a result confirmed under strict out-of-fold evaluation. These results support prefill-guided scheduling as a practical token- and sample-saving mechanism for test-time scaling.
Chat is not available.
Successful Page Load