Why Do LLM Agents Fail to Explore New Environments? A World-Modeling Perspective
Abstract
Large Language Model (LLM) agents often struggle to explore unfamiliar environments whose task-specific dynamics differ from their pretraining experience. In these settings, standard Reinforcement Learning (RL) can improve single-trajectory success while eroding behavioral diversity. We identify exploration collapse, where Pass@k, the probability that at least one of ksampled trajectories succeeds, falls during training even as Pass@1 rises. We attribute this failure to inadequate internal models of environment states and transition dynamics. We introduce SPA, a post-training framework that first collects self-experience trajectories, grounds them with true environment states, and uses supervised finetuning to internalize a world model before policy optimization. Across Sokoban, FrozenLake, Sudoku, and ALFWorld, SPA yields higher reported downstream agentic RL scores in the evaluated settings. On Qwen2.5-1.5B-Instruct, it increases Sokoban Pass@1 from 25.6% to 59.8%. Controlled studies show that the gains depend on grounded state representations, explicit transition modeling, sufficiently exploratory data collection, and adequate data quality and coverage. These results position internal world modeling as a practical route to more robust adaptation of agents in new interactive environments.