When the Loop Closes: Diagnosing and Repairing RLVR for Small Tool-Use Agents
Abstract
Single-turn reinforcement learning with verifiable rewards (RLVR) is the default recipe for post-training small tool-use agents after supervised fine-tuning on distilled trajectories. We report a controlled null result: across seven RLVR objective formulations—spanning token-, sequence-, chunk-, and multi-objective-level ratio construction—sharing one anchor checkpoint, reward stack, and training budget, no formulation improves over its SFT initialization beyond seed noise. Training telemetry reveals why. On-policy importance ratios degenerate to unity, within-group reward variance collapses early on every run, and the dominant evaluation failures—multi-turn trajectory pathologies such as redundant-reload loops that exhaust the turn budget—never occur inside single-turn rollouts, leaving the reward blind to them. Deliberately raising rollout entropy extends gradient-alive windows yet moves outcomes by nothing: the bottleneck is information, not exploration. We close the causal loop by executing the tool loop during rollout, letting the same rewards observe whole trajectories. This repairs the smallest model's collapse outright: the 2B SFT anchor, which completes only 4.3% of evaluation trajectories, reaches 58.7% under multi-turn GRPO (p = 4.5×10⁻¹²⁹, n = 3 seeds), matching a 4B SFT agent at half the parameters; and grounded finals (answers numerically traceable to tool payloads) rise in lockstep, 3.7% → 56.0%. Checkpoint-level analysis traces the gain to the pre-saturation training window: 48 of 52 points land by step 25; post-saturation steps neither help nor preserve the peak. Whether credit is assigned per trajectory or per turn is secondary in this regime (p = .086, consistent direction across seeds): once outcome rewards saturate at structural correctness, per-turn signal is largely redundant and the binding constraint shifts to reward informativeness. Separately, two infrastructure masquerades each moved scores 10–14 points before trace forensics exposed them. We contribute the diagnosis, the repair, and the forensic protocol that made both measurable.