Filter Before Judging: Decision-Scope Evidence Projection for Robust LLM Judging
Abstract
LLM judges increasingly decide whether an agent action is executed, blocked, or escalated, yet a judge often sees an entire trajectory when its decision concerns one typed outbound action, so irrelevant dialogue and other parties’ speech can flip a consequential verdict. We formulate this as a paired causal robustness problem over the judge’s input view and solve it at the prompt layer, where a hosted judge that exposes neither weights nor logits can still be controlled. Decision-Scope Evidence Projection (DEP) deterministically retains the decision contract and the minimal typed action tuple while deleting trajectory-level nuisance fields. Consensus-Selective Judging (CSJ) executes only when full-view and projected- view judgments agree, giving a one-sided safety guarantee and an abstention option. Over 6,440 arm-level decisions from 115 independent privacy-agent scenarios, two closed-weight judges, and an untouched 15-scenario held-out test set, DEP improves accuracy by more than 30 points in every one of the four judge–set cells (all paired 95% intervals exclude zero), mainly by removing false positives, and shortens the prompt by 56%, so CSJ’s two calls cost under 1.5×a single full-view call. Projection alone, however, increases false negatives on the development set. CSJ prevents about 70% of those while executing about three in five decisions, with a false-safe release rate among positive trials of 2.48% for DeepSeek V4 Flash and 2.03% for Qwen 3.8 Flash, and a graded variant traces the coverage–risk curve. Projection is a strong robustness intervention, while abstention and finite-sample certification provide a path toward risk-controlled deployment.