When a Fixed Task Generator Falls Behind: Toward Adaptive Curricula for Training Terminal Agents
Abstract
Terminal agents can be trained through reinforcement learning on tasks with automatically verifiable outcomes, but direct training on challenging tasks often provides little learning signal because the initial policy rarely succeeds. A promising alternative is to construct curricula of intermediate tasks that are easier than the targets while exercising the capabilities required to solve them. Self-play offers a natural mechanism for constructing such curricula: a Generator proposes intermediate tasks, while a Solver learns from them and provides feedback that can guide future generation. Although this approach has shown promise in simpler domains such as mathematical reasoning and single-turn code generation, extending it to long-horizon, interactive terminal environments presents substantial algorithmic and engineering challenges. As a step toward Generator--Solver co-training, we study how far a strong but fixed Generator can go in providing a useful curriculum for terminal-agent training. We use GPT-5 to generate simplified variants of challenging terminal tasks and train Qwen3.5-9B on both the original and generated tasks. We find that training on generated tasks improves Solver learning compared with training only on the original tasks, showing that a fixed Generator can provide useful intermediate training signal. However, our analyses also reveal two limitations. First, the simplified tasks become progressively easier for the Solver as training proceeds, making the fixed curriculum less informative over time. Second, while prompting the Generator for different difficulty levels produces a coarse ordering of tasks, it does not reliably calibrate task difficulty to the Solver’s current capabilities. Together, these observations highlight the limitations of fixed task generation and suggest that adapting the Generator as the Solver improves may provide a more effective curriculum. We view this as motivation for future work on adaptive Generator--Solver co-training in terminal environments, in which the source of training experience keeps learning alongside the agent it trains.