EvoBelief: Falsifiable World Models for Self-Improving LLM Agents
Abstract
An LLM agent ships with a frozen prior, everything its weights already believe about how the world works. The environment it is dropped into has its own dynamics, some absent from pretraining, some contradicting it, and some that change while the agent is running. An agent can be procedurally competent and still fail there, because what it misunderstands is not the procedure but how this environment responds to it. Existing self-evolving agents mostly improve how the agent acts, by accumulating trajectories, reflections, and skills, while methods that do evolve a world model improve predictions that rarely become better actions, a failure known as the prediction-to-action gap. We introduce \textbf{EvoBelief}, which makes the agent's belief about environment dynamics the object of self-evolution and runs the scientific method on the deployed environment. The agent states what it believes as falsifiable hypotheses, tests them with controlled experiments that change one factor at a time in a sandbox forked from the live state, and lets evidence, and only evidence, promote a hypothesis into a rule that may change behavior. When the environment drifts, counterexamples revoke that right. All adaptation lives in this external belief state, so the language model is never updated, and every changed action stays attributable to the rule, the experiments supporting it, and the decision it flipped. Across four benchmarks spanning static, drifting, and scientific-discovery regimes, EvoBelief learns genuine environment dynamics, converts them into the highest task success under both actor settings, and leaves a belief state that can be opened and read.