When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenges
Abstract
Automated terminal-agent training relies on generated environments and verifiers to turn open-ended interaction into reward. Yet a runnable Docker image and executable test suite do not guarantee that the resulting reward is faithful, informative, or even attributable to agent behavior. We present a pipeline that generates containerized terminal tasks and their executable verifiers, and use it to demystify failure modes in task generation and live-environment evaluation. Our analysis separates four requirements that are often conflated: structural validity, verifier faithfulness, policy solvability, and infrastructure reliability. On a roughly 10\%-solvable task distribution, PPO training of Qwen2.5-3B for 228 steps does not improve on its 7.8\% base pass@1, consistent with predominantly zero reward. For a GRPO training of Qwen3.5-9B, it overfits on a relatively easy dataset within 20 steps. Conversely, in a controlled Qwen3.5-9B comparison, adding hard tasks lowers mean rollout pass@2 from 81.3\% to 20.6\%, restoring informative failures without changing the training configuration. We further identify near-duplicate task generation, incomplete anti-shortcut safeguards, Dockerfile failures, asynchronous pass@k bias, and concurrency-induced environment errors as distinct threats to measurement. These findings should motivate the community for a terminal agent specific evaluation harness that pairs structural checks and rollout-derived solvability with extensive verifier audits.