Diagnose, Then Select: What Makes Critique-Based Refinement Effective?
Hyun Ryu ⋅ Jihwan Oh ⋅ Sumyeong Ahn ⋅ Se-Young Yun
Abstract
Large language models increasingly rely on critique to refine their reasoning, yet existing approaches treat critique as an undifferentiated block of feedback and provide more information than may be necessary. We investigate which parts of a critique make an error repairable by decomposing one structured record into isolated correctness-verdict, error-location, and error-diagnosis views, including a full critique. Across three refinement protocols, two model families, and two mathematical-reasoning benchmarks, diagnosis-only feedback recovers $98.8\%$ of the wrong-to-correct rate of a full critique while using $30.9\%$ of its feedback tokens; verdicts and locations recover only $42.6\%$ and $62.3\%$. Correction alone, however, is an incomplete objective: critique also destabilizes answers that were already correct. Under solver--critic refinement on MATH-500, diagnosis corrects $8.8\%$ of the weaker Llama solver's wrong answers but corrupts $25.7\%$ of its correct ones, a net loss of $7.4$ points, while the same feedback gains $3.8$ points for the stronger Qwen solver. We formalize this base-rate dependence through a break-even accuracy $\alpha^*$ and propose **DiSel**, which gates diagnosis-conditioned revisions by confidence. Gating substantially reduces corruption, but its benefit is model- and benchmark-dependent. Our results identify two bottlenecks in refinement: generating the causal information needed for repair and reliably deciding whether to accept it.
Chat is not available.
Successful Page Load