Position: Clinical Trust in Health Agents Should Rest on Evidence Paths, Not Answer Checks
Abstract
Evaluations of health agents increasingly certify trustworthiness by answer checking: if every number, identifier, and status in the generated text traces to a retrieved source, the output is treated as grounded. We argue that this standard cannot certify clinical trust, because a system can produce a conclusion that is externally defensible while remaining unsupported by its own retrieval. Our evidence is an adversarial audit of a drug-target prioritization agent, a system whose recommendations sit at the head of every therapeutic pipeline. The agent records tool outputs in an append-only ledger, constructs evidence tables in code, and uses a language model only to route tools and to generate a verdict with a short rationale. A seven-week hardening effort eliminated numerical drift, yet the audit confirmed 18 distinct defects, none detected by the validator: the model asserted relationships absent from the retrieval, importing them from pretrained knowledge. A longitudinal IL6R-coronary-heart-disease example shows keyword-triggered checks responding to wording while the unsupported inference persisted, and later phase-3 evidence in the same pathway shows why a conclusion that was defensible at the time must not be mistaken for one that was established. We also find that a pass rate alone can reward outputs that make few checkable claims. We therefore take the position that health-agent evaluation must verify an evidence path for every clinical conclusion, require unretrieved links to be retrieved, declared as assumptions, or abstained from, and report claim coverage beside pass rates.