Characterizing Diverse Failures in Agentic Systems with the OURI Testbench
Abstract
Modern agentic AI systems combine large language models (LLMs), software tools, and increasingly complex orchestration strategies, making their behavior difficult to characterize from aggregate benchmark scores alone. Failures are often diverse, context-dependent, and distributed across trajectories rather than summarized by a single accuracy number. We present OURI, a testbench for Observing, Understanding, Recommending, and Improving Agentic Systems. OURI evaluates combinations of open-weight LLMs, orchestration strategies, tools, and task datasets while preserving observable execution traces. LLM-powered analysis agents compress those traces into structured measurements that can be aggregated into configuration characterizations and later used as behavioral descriptors for diversity-aware failure analysis. Selected demonstrations show that performance depends jointly on model, task, and orchestration, and that likely failure mechanisms differ across configurations; appendix views complement this picture with confidence when wrong and prompt composition by source. The resulting characterizations support recommendation and iterative improvement, and sketch a path toward diversity-driven search over failure regimes.