Not All Low-Confidence Tokens Are Equal: Calibrated Confidence for Efficient Test-Time Reasoning
Abstract
Test-time scaling boosts the reasoning performance of large language models (LLMs) on challenging tasks but incurs substantial computational overhead. The existing training-free approaches adopt token-level confidence as a proxy and terminate the trajectories once the confidence falls below a threshold. However, these methods view all low-confidence tokens uniformly, overlooking the fact that some of them actually reflect benign generations. We introduce Calibrated Confidence (CalConf), a training-free proxy for test-time scaling in LLM reasoning. The design rests on a simple empirical observation: flawed trajectories tend to be long. CalConf therefore learns a length-calibrated token-weight map and uses it to selectively terminate degrading trajectories during generation. We theoretically prove that CalConf attains the desired early-stop rate in practice without per model tuning. Across five benchmarks (on mathematical reasoning, scientific reasoning, and code generation tasks) and three open-source LLMs, CalConf consistently matches or exceeds offline majority-vote accuracy while reducing generated tokens by 20–44%. Token-level analyses trace these gains directly to CalConf’s ability to isolate harmful patterns that uniform-confidence baselines cannot distinguish.