TraceJudgeBench: A Unified Benchmark for Validating LLM-as-a-Judge on Agent Execution Traces
Abstract
LLM-as-a-judge methods are increasingly used to evaluate LLM-based agents, yet the benchmarks used to validate these judges are fragmented: they define incompatible taxonomies, annotation schemes, and error distributions that make cross-study comparison unreliable. We introduce TraceJudgeBench, a unified benchmark that preserves native trace representations from ten existing judge-validation datasets while harmonizing annotations under a shared two-level taxonomy. The fine-grained L2 taxonomy covers 10 categories grouped into 5 coarser L1 classes. We evaluate a prompted LLM, TRAIL, AEGIS, and a decomposed ensemble judge using the same backend and taxonomy-aligned metrics. Results reveal strong category-level variation and a consistent false-error bias on clean traces, with rates from 34% to 78% across the three established judges and reaching 100% for a decomposed ensemble run without abstention. The prompted baseline stays between 67% and 78% across three model families, so the effect follows the judge design across backbones. Splitting the clean traces by length shows that this average hides most of the structure: every judge degrades a further 37 to 50 points on long traces, an effect a length-concentrated corpus could not expose. These findings demonstrate that aggregate scores mask critical judge weaknesses and that clean-trace recognition must be an explicit evaluation target. Benchmark materials are available at https://anonymous-hf.up.railway.app/a/rj6gqe9zn2xp/. The accompanying code is available at https://anonymous.4open.science/r/TraceJudgeBench/.