Where Does the Error Live? A Three-Level Framework for Evaluating Multi-Agent Workflows in Production
Ahmed Elmokashfi ⋅ Stefano Costanzo
Abstract
Agents that recommend under uncertainty---in medical triage, incident response, security operations---produce outputs whose correctness is revealed only when a human resolves the case. We present a three-level evaluation framework that exploits this structure: resolution records, generated as a byproduct of normal operations, serve as free ground truth requiring no labeling effort. Deployed on a large infrastructure diagnostic workflow, which comprises more than 90 steps and 70 individual agents, processing hundreds of incidents weekly, the framework operates at three granularities. Level1 scores recommendations against resolutions and, firing on every case closure, turns quality into a time series that distinguishes interventions extending system reach from those sharpening specificity. Level2 localizes failures within the agent graph via causal-graph alignment: an anti-anchored procedure reconstructs the resolution's causal chain independently and traces misalignment to an originating agent, classifying errors into Knowledge, Reach, Attention, and Plumbing; Attention (evidence available but unused) accounts for 68% of failures, directing remediation to synthesis logic rather than knowledge or tooling. Level3 verifies reasoning trajectories against safety rules before the system acts, using a deterministic classifier (F1 0.853, $\pm$0.005 variance) that sidesteps the LLM extraction bottleneck ($\pm$0.10 F1 across draws). We characterize reliability at each level: agreement degrades from $\kappa$=0.631 on coarse verdicts to 0.431 on failure class, establishing that evaluation instruments require the same stability engineering as the systems they assess.
Chat is not available.
Successful Page Load