When Has an Agent Seen Enough? CASE for Evidence-Aware Scientific Autonomy
Aisha Opaluwa ⋅ Brandone Fonya
Abstract
When Has an Agent Seen Enough? CASE for Evidence-Aware Scientific Autonomy Scientific AI agents increasingly complete experimental workflows end-to-end, yet completing a workflow is not the same as reasoning correctly about the evidence within it. A recent large-scale evaluation of scientific agents across eight domains found that evidence is ignored in 68% of reasoning traces and refutation-driven belief revision occurs in only 26%, even when task outputs appear successful [1]. This suggests that current autonomy is calibrated to task completion rather than to the evidence that should justify it. Recent work has begun to address this gap: sequential agentic abstention formalizes when an agent should act, gather more evidence, or stop [2], and explicit belief-state representations have been proposed to guide autonomous scientific discovery [3]. However, these approaches are evaluated in synthetic or general-purpose settings, leaving open how evidence-calibrated autonomy behaves when evidence comes from real, costly scientific experimentation. We introduce CASE (Calibrated Autonomy for Scientific Experimentation), an agent framework working toward a machine learned decision policy for evidence-aware decision-making in real computational science. The current system maintains an inspectable record of competing hypotheses and classifies each new result as sufficient to conclude, worth rechecking, or genuinely inconclusive, distinguishing recurring failure patterns from one-off ambiguous outcomes rather than re-diagnosing every case independently, and serves as the interpretable baseline this policy will be learned against. We are developing CASE and evaluating it using density functional theory simulations of binary (Mn,X)WO$_4$ tungstates, replaying completed Quantum ESPRESSO traces sequentially so the agent receives evidence incrementally rather than the final outcome in advance. In an initial run across four calibration materials, CASE separated two cases whose energy gap fell inside the noise floor but was judged worth a denser-mesh recheck from one case whose gap sat at the noise floor's own resolution limit and was reported as genuinely degenerate; a naive baseline lacking this distinction reported all three identically as a flat escalation. For every rerun, recheck, or new-candidate decision, CASE generates a corresponding Quantum ESPRESSO input file as an auditable record of what evidence justified the decision, including, for a new candidate material with no prior run, a prediction made by matching it to a calibrated case that shares the relevant underlying physical property, explicitly flagged as pending confirmation rather than settled. This setup lets us ask: (1) does persistent evidence tracking reduce unnecessary reruns and improve recognition of ambiguous outcomes compared to independent, per-result reasoning? (2) does accumulated calibration knowledge improve decisions on unseen simulation cases without encouraging unsupported extrapolation from insufficient precedent? We anticipate that grounding evidence-calibrated autonomy in real experimental costs and failure modes, rather than synthetic benchmarks, will clarify whether current approaches to scientific agent autonomy generalize beyond the settings in which they were developed, and offer a concrete testbed for evaluating when scientific AI systems can be trusted to act without human review. References [1] M. Ríos-García, N. Alampara, C. Gupta, et al. AI Scientists Produce Results Without Reasoning Scientifically. arXiv:2604.18805, 2026. [2] H. Luo, B. Wen, L. L. Wang. Agentic Abstention: Do Agents Know When to Stop Instead of Act? arXiv:2606.28733, 2026. [3] X. Wu, S. Yu, Q. Xu, S. Yin. BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery. arXiv:2606.30335, 2026.
Chat is not available.
Successful Page Load