Hard Enough to Learn: Evolving Verified Tool-Use Environments with Code Agents
Timur Ionov ⋅ Nikita Andriianov ⋅ Anastasiia Fedorova ⋅ Maksim Savkin ⋅ Andrey V Galichin ⋅ Vasily Konovalov ⋅ Valentin Malykh ⋅ Daria Pugacheva
Abstract
Large language model (LLM) agents have emerged as a dominant paradigm in artificial intelligence, with their capabilities increasingly determined by their ability to invoke and orchestrate external tools. Reinforcement learning acquires that ability at scale, and it needs verified environments with evolving difficulty to keep the training signal rich. We build a factory that produces verified RLVR data with code agents. An agent authors a domain, tools with implementations, a seeded database, and declarative task templates. Deterministic code then instantiates each task, derives its verified gold tool-call chain. One deterministic predicate validates the gold, measures each task's difficulty, and serves as the reward, so no judge sits in the loop. The factory yields 64 domains and 6,989 verified tasks, works with any code agent han and 19 independent rebuilds of the same specifications by a second coding agent. Per-task difficulty estimates calibrate the corpus to a learnable band and let the code agents harden the saturated tasks. The band-calibrated corpus beats the same tasks drawn at random and beats twice as many drawn at random. Training on the band-calibrated corpus improves Qwen3.5 across scales on five held-out evaluation benchmarks: WorkBench, Workplace Assistant, GAIA2, BFCL multi-turn, and $\tau^3$-bench. An ablation traces the gains to the corpus rather than the optimizer, and supervised fine-tuning on it installs tool use in weak non-Qwen bases.
Chat is not available.
Successful Page Load