Stop or Restart? Principled Inference Control for Large Reasoning Models via the Pandora's Box
Xiaosong Yuan ⋅ Xiaofeng Zhang ⋅ Yijia Zhang ⋅ Renchu Guan ⋅ Ying Wang
Abstract
While large reasoning models (LRMs) can solve complex tasks via generating extended Chain-of-Thought traces (Long-CoT), longer traces often increase confidence without increasing correctness. In this work, we formalize such failure via a \emph{self-conditioned evidence} model: generated tokens provide information about the model's current working hypothesis rather than the ground-truth answer. This yields a simple diagnostic principle: entropy, fluency, and answer stability can certify within-trajectory self-consistency while remaining causally disconnected from correctness. Building on these findings, we propose \textbf{Faithful Pandora Inference (FPI)}, a minimal training-free controller that keeps the base LRM frozen and fits only a small calibration map on held-out data. At each checkpoint, FPI estimates calibrated correctness $r_t$ and the marginal evidence gain $m_t$ of continuing the current trajectory. It stops only when $r_t$ reaches a user-specified reliability target; if $r_t$ remains insufficient and $m_t$ falls below cost, it reallocates the remaining budget to a diversified restart. Across six reasoning benchmarks and multiple DeepSeek-R1 distilled models, FPI improves the accuracy-efficiency frontier, reduces expected calibration error, and improves selective accuracy on high-confidence outputs. Ablations on confidently wrong trajectories show that calibrated stopping and marginal-gain restarts address different failure modes and that the gains are not explained by additional samples alone.
Chat is not available.
Successful Page Load