Hierarchical Experimentalist Agents
Abstract
Large language models (LLMs) are increasingly used to act in the real world and augment human decision-making, yet most systems rely on parametric knowledge acquired through imitation, optionally augmented by post-training, retrieval, or search over fixed data. This paradigm struggles in novel domains and on complex queries that require information unavailable from prior knowledge alone; knowing the laws of physics, for example, does not by itself enable reliable reasoning or long-horizon action in a complex physical system. We argue that agents therefore need \textit{active experimentation}: the ability to gather targeted query-specific evidence, discover general principles of unseen environments, and convert interaction experience into reusable skills. We introduce \textbf{Hierarchical Experimentalist Agents (HExA)}, an in-context, experiment-centric self-improvement framework that (1) iteratively designs and refines query-relevant experiments, (2) incrementally distills experience into reusable and composable skills that accelerate experimentation within and across tasks, and (3) integrates experimental evidence to act or answer queries. HExA is entirely in-context and training-free, works with black-box models, and requires no external supervision, oracle solutions, or offline data. To evaluate active experimentation, we introduce \textsc{Interphyre}, built on the PHYRE 2D procedural physics environment with tool-calling and intervention APIs for proposing and testing hypotheses. Current LLM agents struggle substantially in this setting: on the hardest \textsc{Interphyre} level, Claude Sonnet 4.6 achieves only 2\% success, while HExA reaches up to 77\%. We observe similar gains with open-weight models and over agentic baselines including ReAct and Reflexion. Moreover, using only skills transferred from easier levels and \emph{no target-level experimentation}, HExA achieves 44\% success, demonstrating that its learned skills are reusable and generalize across tasks. These results show that learning through experimentation can help agents discover new knowledge and efficiently make progress on novel tasks.