Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations
Xinshun Feng ⋅ Ziqi Miao ⋅ Lijun Li ⋅ Jing Shao
Abstract
Large language model (LLM) agents are increasingly deployed in scientific research, where reliability is critical, and the underlying knowledge is densely interconnected. In such settings, hallucinations are particularly damaging: a single erroneous claim on a foundational concept can propagate through multi-step reasoning and corrupt entire trajectories. Existing hallucination benchmarks largely operate at the surface level, treating facts in isolation and relying on uniform accuracy metrics that ignore this topological structure. We address this gap with \textsc{Schema}, the first evidence-grounded, topology-aware evaluation framework for scientific agents. Built on an automated concept-graph construction pipeline, \textsc{Schema} provides two complementary diagnostic instruments. A trajectory hallucination pipeline audits intermediate reasoning at scale via a topology-weighted severity score ($\mathrm{HS}^w$), while a multi-agent counterfactual attribution module pinpoints the causal mechanism behind selected failures. The framework supports diverse task formats spanning multi-hop reasoning, claim verification, and programmatic experimental execution. Across two biomedical subdomains and eleven contemporary LLM agents, \textsc{Schema} reveals that hallucinations concentrate at a small set of highly connected knowledge hubs, and that final-answer accuracy decouples from trajectory honesty—models often reach correct conclusions through structurally flawed reasoning. These results indicate that for high-stakes scientific applications, terminal accuracy alone is an insufficient signal of agent reliability, motivating mechanism-level evaluation grounded in knowledge topology\footnote{Code is available at \url{https://anonymous.4open.science/r/SCHEMA_NIPS2026}}.
Chat is not available.
Successful Page Load