Do Agent Guardrails Explain Themselves? Causal Faithfulness of Tool-Call Gating Explanations
Abstract
When an AI agent's tool call is allowed or blocked by a policy gateway, an LLM typically generates the natural-language explanation shown to users and auditors. We study this as a data-to-text generation task with an unusual property. The input includes an executable decision function, so the correct content of the explanation (the minimal sufficient reason for the decision) is exactly computable. This yields a reference-free, fully automatic evaluation of content selection in generated explanations, via causal precision (is cited content actually part of a true reason?) and causal recall (is some complete reason covered?). Across four LLMs, end-to-end generation produces fluent explanations whose selected content rarely contradicts the input record yet is nonetheless about half irrelevant, at every model scale we test. The failure is content selection, not surface realization. Constraining length does not fix it. Under a one-sentence budget, models keep distractors and drop true reasons. Separating content determination from surface realization (an oracle selects content, the LLM only realizes it) restores near-perfect content precision and recall at equal fluency, and the guarantee survives audience adaptation into end-user and executive registers at a small realization cost. We release the benchmark, oracle, and evaluation suite.