LUCA: A Synthetic World for Experimental Reasoning
Abstract
An experimental agent can succeed for the wrong reasons. Yet most evaluations observe the outcome, not the experiments, evidence, and inferences that produced it. We introduce LUCA, a configurable synthetic world for evaluating this process under fully known ground truth. Agents learn unknown regulatory systems through budgeted experimentation, then predict unseen interventions, diagnose hidden faults, and attempt rescues. This lets us evaluate experiment selection, prediction, diagnosis, and intervention separately. Across seven frontier LLMs and three baselines, these capabilities dissociate sharply. Agents frequently rescue malfunctioning systems while diagnosing the wrong fault; diagnosis varies 7× across models while prediction varies only 1.4×. Models also adopt surprisingly different experimental strategies, and greater exploration does not reliably yield better performance. In one failure mode, agents often perform the experiment needed to distinguish competing causes, yet fail to use the resulting evidence. LUCA turns these hidden failures into measurable behaviours, providing a controlled environment for studying where experimental agents succeed and fail.