Coding Agents as Autonomous World-Model Researchers
Abstract
Can coding agents make useful machine-learning research progress when the direction of improvement is left open-ended? We study this question in world modeling, asking whether coding agents can autonomously improve the predictive accuracy and long-horizon rollout quality of a given world model under a fixed compute budget. Each session provides an agent with a base world model, structured trajectories, and a read-only evaluator. The agent forms its own hypotheses, edits the model or training code, evaluates the result, and decides what to try next. We assess this setting across eight games and four base architectures. To distinguish genuine model improvement from overfitting to the validation signal, we evaluate the discovered models on a separate, human-designed suite of scenarios that is never exposed during the search process. We also compare against compute-matched random search and Optuna/TPE, repeat the experiments across additional seeds, and audit the sessions for scorer tampering. Importantly, the interventions are proposed, implemented, and selected entirely by the agents, making the resulting improvements a product of the autonomous model-development loop rather than human-specified optimization. We run 64 independent agent sessions, each lasting six hours. Across these runs, Codex (GPT-5.4) and Claude Code (Opus 4.6) improve the base model on the held-out test split in 63 cases, with a mean improvement of +0.196. The gains are largest at longer rollout horizons, where autoregressive error compounds: +0.205 at h=10 and +0.215 at h=20, compared with +0.056 at h=1. In 58 of the 64 runs, the best-performing intervention changes the architecture, objective, representation, rollout procedure, or inference rule rather than simply tuning a scalar hyperparameter. The resulting models mostly improve on the human-designed test scenarios, showing that the discovered interventions generalize beyond the validation signal used during search. These results provide evidence that current coding agents can carry out a constrained but end-to-end machine-learning research process.