GlobeBench: Evaluating Text-Based World Model Capabilities for Scalable Environment Simulations
Abstract
Text world models offer a way to reduce dependence on expensive external environments during agent training, but plausible observations do not necessarily preserve state, apply the correct transition, or enforce an environment's rules. We introduce GlobeBench, a 1,632-item benchmark for text-based world simulation spanning five domains, four capabilities, and 20 fine-grained diagnostics (19 primary construction skills and one cross-cutting label). Each item asks a model to predict an execution-derived next state from a multi-turn interaction history and a final action. We evaluate ten general-purpose language models and two models trained specifically for world simulation using deterministic checks, followed by blinded semantic adjudication when structural comparison is insufficient. The best model achieves 28.13% task success, and neither simulation-tuned checkpoint outperforms the strongest general-purpose models. Across the ten general-purpose models, 55.6% of 13,266 failed decisions contain at least one incorrect concrete value, often inside an otherwise plausible response. Success falls from 43.15% for local effects to 3.66% for the broadest state changes, and drops below 11% when the relevant state-setting event is more than 40 turns away. These results show that interface plausibility and aggregate accuracy conceal distinct failures in state tracking, rule enforcement, and side-effect propagation.