Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
Abstract
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25\% at matched accuracy across Gemma, Qwen, and Nemotron models on mathematical, scientific, and coding reasoning benchmarks, with accuracy--efficiency tradeoffs comparable to methods that explicitly optimize for shorter reasoning. Moreover, on Gemma, applying confidence supervision on top of length-penalty RL yields further token reductions, suggesting that the two forms of training can be complementary. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized.