SteadyThought: Mitigating LLM Under-Thinking via Thought-Level Preference Optimization
Abstract
Flexible switching between reasoning trajectories (i.e., thoughts switching) has significantly enhanced the reasoning capabilities of Large Reasoning Models (LRMs). However, existing models often switch excessively yet fail to sustain promising reasoning thoughts---a phenomenon termed ''under-thinking''. While recent efforts suppress switching to mitigate this, such over-correction may discard valuable trajectories. To address this challenge, we propose Steady Thought (ST), a novel thought-level preference optimization framework. ST formalizes under-thinking as a preference issue at switching points, forcing single-path continuations from identified points to yield high-quality reasoning. Then, ST performs thought-level preference optimization by treating the newly generated response as preferred and the original one as dis-preferred. Experiments across multiple models and datasets show that ST effectively reduces token consumption while maintaining or even improving accuracy. It reduces output length by up to 63.8% while improving accuracy by up to 11.4%. Further analysis suggests that ST helps models acquire a generalizable ability to reduce redundant switching across domains and languages.