Can Compute, Won't Execute: Decision-Boundary Probing of Policy Faithfulness in Tool-Using Agents
Abstract
Tool-using language model agents can retrieve a written policy, access every relevant financial field, and still execute a decision rule that ignores the policy. We introduce LedgerTwin, a benchmark of 2,400 executable ledger worlds organized into 600 four-variant tuples, each a fully functional ERP snapshot with a four-clause credit-hold policy and controlled counterfactual variants that disentangle policy faithfulness from surface accuracy. We also introduce CLADD (Clause-Level Auditing via Decision Decomposition), a behavioral diagnostic that sweeps a controlled variable across an agent’s decision boundary to recover which of the 16 possible clause subsets best explains its behavior, without access to model internals. Across six evaluated models, we identify three distinct failure modes: stable simplification (Qwen3-14B consistently executes the invoices-only heuristic despite retrieving the full policy), scale inversion (Qwen3-32B fails to form stable decision boundaries despite its larger scale), and incoherent partial composition (GPT-5.4 applies policy clauses inconsistently, yielding 77.5% non-monotonic decision curves). Exact recovery of the prescribed policy is 0% for all Qwen models and only 2.5% for GPT-5.4. For the two models we probe directly, these failures need not reflect missing capability: Qwen3-14B produces a value nearest to the correct threshold in 82.5% of isolated arithmetic probes, while a 150-token reasoning scaffold increases GPT-5.4’s exact policy recovery from 2.5% to 95.0%. LedgerTwin and CLADD reveal execution failures that aggregate accuracy can conceal, providing a practical diagnostic for evaluating policy-faithful tool-using agents before deployment.