Verifiers Are All You Need
Abstract
The hardest problem in scaling RL today is environment design that feeds production training pipelines with a steady stream of high-quality tasks that stay discriminative as models improve. Static environments built from one-time human-authored tasks go stale fast and become depreciating assets. We design and prove how a continual learning harness, with an asynchronous human-in-the-loop premise, not only scales RL by reducing task authoring time by an estimated 8×, but also proves effective at continual and dynamic calibration. We divide the task creation process into well-defined atomic steps, and gate each step with an expert-calibrated verifier. The harness concentrates human expert time on low-confidence judgements. Clean runs seed the next round. Under these gates an autonomous loop can explore widely, because everything it admits is clean. We prove the design on a hard domain. The benchmark's mainframe suite is built on the decommissioned billing system of a large telecommunications company, a 432k LoC COBOL, JCL and DB2 estate. It runs off the mainframe on a toolchain hardened against the system's recorded behavior. The harness explored 234 task ideas, of which 64 cleared the pre-run design gates. We construct a benchmark of 27 of these 64 tasks, and report the findings in this paper. Each task was admitted at a bar of four execution arms, and the grader's discrimination was measured before any agent ran. Two frontier agents, kimi-k3 under Terminus 2 and Claude Fable 5 under Claude Code, pass 18.5% and 23.0% of trials. Localization is not the bottleneck: in five failures out of six the agent reaches the graded programs, and a task spans several programs, clone legs and archived generations. The agents fail for three reasons: the fix does not cascade through the entire chain, it fills a missing business rule with an unchecked assumption, and the agent verifies it against a compile or a fixture written from the same assumption, which cannot fail. Post-run CI evaluates each rollout on seven dimensions and seven task-quality flags, and qualifies post-training-eligible rollouts. Each qualified rollout is used for behavioral analysis at 5% human grading, both as a model training signal and to feed the task authoring loop.