Scaling Verifiable Environments for Long-horizon Working Agents
Abstract
Work agents operate over digital artifacts to automate multi-step professional knowledge work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents cross-domain scaling despite offering authentic workflows, whereas automated synthesis methods remain limited in environmental complexity, realism, and verifiability. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files, organizes them into an agent-accessible workspace, and examines the workspace to determine which tasks and outcomes it can support. Based on this evidence, WorkForge generates the task and solution plan together with programmatic checks and semantic rubrics, then runs a solver agent to produce a scored long-horizon training trajectory. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 distinct file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance, consistently outperforming existing strong competitors. Controlled analyses further reveal consistent scaling behaviors across both data volume and interaction horizons.