Sampling Is Not Curiosity: Why LLM Agents Should Investigate
Abstract
LLM agents are increasingly used for open-ended tasks, very common in research scenarios where they are required to explore, investigate, or act curiously. At inference time, LLMs sample from a distribution that is fixed by its pre-trained weights and the provided context. Mechanisms like temperature, best-of-N, tree search, or RAG vary how the distribution is drawn, but do not update the agent's model about the environment. We argue that current LLM agents do not explore curiously; they explore stochastically. From how curiosity is understood in the RL literature, it requires a predictive structure over how the environment responds to the agent's actions, which improves online from current interaction, drives action selection through its predicted improvement, and persists in structured form across the episode. Additionally, in its open-ended form, the agent should intrinsically originate what to investigate. Thus, we propose a framework defining curiosity in LLM agents, reserving the term for systems that satisfy it. In doing so, we identify the research directions its implementation opens up.