OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine
Eray Turkel ⋅ Mengsha Sun ⋅ Kartik Ayyar ⋅ Hsiang-Shun Shih ⋅ Xin Wang ⋅ Tiantian Zhang ⋅ Sean Dunigan ⋅ Jack Lu ⋅ Vlad Shcherban
Abstract
We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game engine sessions and scores them with executable checks at edit time and play time simulation. Its eight-tool action space separates observation tools from editing tools, so exploration is explicitly measurable. Most agentic coding benchmarks require exploration, but typically evaluate only final task success, leaving the information contained in exploration traces unused. A live game engine makes that investigation both harder and observable. We measure both pass rates and exploration behavior of $13$ frontier models on $84$ human-curated core tasks, each attempted $16$ times. The current generation of models finds our task set difficult: the best model solves $51.7\%$ of tasks on a single attempt and $39.4\%$ five times out of five. Six tasks are unsolvable by any model. Models at the frontier reach similar pass rates by solving different tasks. The top five sit within $1.1$ pp of each other yet share only $27$ of the $58$ core tasks they solve between them. Splitting tasks by the kind of work they require separates the top five by $5.0$pp on script-authoring and $12.5$pp on scene-change tasks. We show exploration behavior carries signal about whether a run succeeds. Holding task and model fixed, reaching the objects a reference solution touches before acting on them is associated with $+13.4$pp increase in likelihood of success on scene-change tasks and $+9.8$pp on script-authoring tasks. The task suite, the place files it runs against, the per-task annotations, the plugin to run the tasks inside Roblox Studio, and an updated leaderboard are available under the MIT license in a public repository.
Chat is not available.
Successful Page Load