Why Do LLM Agents Fail to Explore New Environments? A World-Modeling Perspective
Abstract
Pretrained large language models encode broad semantic knowledge, yet their representations can remain difficult to use for sequential decision-making in unfamiliar interactive environments. We study this gap through agentic reinforcement learning and identify exploration collapse: as policy optimization proceeds, Pass@1 can improve while Pass@k declines, indicating that the agent concentrates on a narrow behavior instead of discovering alternative successful trajectories. We attribute this failure to weak grounding of environment states and action-conditioned dynamics. We introduce SPA (Self exPerience Agent), a post-training framework that makes pretrained representations more actionable before policy optimization. SPA lets the base model collect its own interaction trajectories, replaces its state beliefs with structured ground-truth current and successor states, and uses supervised finetuning to internalize transition dynamics. The resulting model initializes downstream reinforcement learning. Across Sokoban, FrozenLake, Sudoku, and ALFWorld, SPA improves the reported downstream agentic RL scores in the evaluated settings. On Qwen2.5-1.5B-Instruct, Sokoban Pass@1 rises from 25.6% to 59.8%. Ablations show that correct state abstractions, explicit transition supervision, policy-guided data collection, and sufficient coverage are all important. These results show how grounded predictive representations can bridge pretrained knowledge and reliable action under environment shift.