Against Epistemic Shortcuts: Evidence-Grounded Verification for LLM Agents
Abstract
Task completion is a commonly used evaluation metric for LLM-based agents. However, even within a trajectory deemed successful, not every action is necessarily evidence-grounded or self-contained. Such cases may arise when an LLM hallucinates or fails to verify important information yet still succeeds by chance, or when human demonstrators with strong environment-specific priors rely on internal knowledge and omit steps that would otherwise be necessary. In this paper, we refer to such behaviors as epistemic shortcuts, which can cause direct failures at test time or create confusion during agent learning when they appear in training data. To address this issue, we propose Evidence-Grounded Verification for Interactive LLM Agents (EVIA), featuring an epistemic monitor that checks action-level evidence sufficiency by identifying the factual assumptions required by a proposed action and checking whether they are supported by the observed trajectory. It supports both test-time control, by redirecting unsupported actions to alternative actions, and training-time rectification, by converting pseudo-successful trajectories into evidence-grounded demonstrations. Our evaluation demonstrates that EVIA outperforms baselines across multiple LLM-based agent domains.