What Must You Log? Decision-Sufficient Trace Signals for Agent Failure Attribution
Abstract
Deciding whether an autonomous agent failed, and which step caused it, is a decision made entirely from what the system recorded, and there is no agreement on what that should be. Across ten public agent-failure corpora median trace size spans a 163-fold range, only four separate the model input from the model output, and no published evidence supports either choice. Holding TRAIL's taxonomy, judge prompt and scoring code fixed, we vary only which trace fields the judge receives. Failure localisation needs one content field, what the agent itself wrote: withholding it collapses accuracy from 0.39 to 0.06 on one judge and from 0.37 to 0.11 on a second, while retaining it alongside step identifiers and two near-free metadata fields recovers full-trace accuracy at 2.33\% of the raw token volume. The replayed model context is 84\% of trace content and the field where prompt data and retrieved documents reside. It never helps localisation, and supplying it whole costs both judges 0.13 location accuracy; 85.7\% of it repeats text already recorded in the same trace. The retention rule that follows is to record what each agent wrote and the identity of the step, and to treat context replay and ordering as optional. It is scoped to that decision: the benchmark's categorisation metric peaks on the raw dump instead, so a deployment that must also categorise should measure the rule for that decision rather than inherit it.