Self-Trained Verification for Training- and Test-Time Self-Improvement
Abstract
Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate over rounds while accuracy stagnates, and when feedback is too generic to act on; self-training fails similarly when bad self-generated data are added to training. Better verifiers would unlock both, but the capability we want to train, i.e., catching errors the generator cannot detect in its own work, has neither finetuning data nor a verifiable reward. To address this challenge, we propose self-trained verification (STV). Our key observation is that, while a model cannot catch these errors alone, it can when shown the reference solution. We turn this asymmetry into a supervision target and train the verifier to imitate a more informed version of itself. At test time, STV substantially improves V-R loops on hard problems, while standard alternatives (e.g., SFT, RL on verifier scores, and even meta-verifiers) do not. STV roughly doubles accuracy on hard DAPO math and lifts it 14x on the hardest SciKnowEval problems (1.5% to 21%). At training time, starting from an RL-converged generator, putting the STV verifier in the loop yields a further 33% relative gain in test-time pass@1. More notably, the generator's standalone pass@1, with no verifier at inference time, climbs 30% relative past where continued RL had converged. Hence, the next frontier in reasoning may lie in how we train verifiers.