Tool Verification for Test-Time Reinforcement Learning
Abstract
Test-time reinforcement learning (TTRL) has emerged as a promising paradigm for adapting large reasoning models (LRMs) on unlabeled test inputs, using self-consensus rewards derived from majority voting over sampled rollouts. However, majority voting can mistake popularity for correctness: a spurious yet high-frequency unverified consensus may become a biased reward signal, causing test-time RL to reinforce frequent but wrong answers and collapse into an incorrect mode. We address this false-popular failure mode with T^3RL (Tool-Verification for Test-Time Reinforcement Learning), a verification-aware test-time RL framework. T^3RL grounds pseudo-label construction in external tool evidence. Concretely, a verifier uses external tool evidence, such as code execution, to upweight verified rollouts during verification-aware voting, producing more reliable pseudo-labels for training. Across mathematical reasoning benchmarks of varying difficulty, including MATH-500, AMC, and AIME 2024, and diverse backbone families, T^3RL significantly improves over TTRL, with stronger performance on harder problems. More broadly, T^3RL is positioned as a verified online data synthesizer, highlighting the role of test-time tool verification in reliable online adaptation.