G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm
Zijun Yu ⋅ Yu Gu ⋅ Vahid Partovi Nia ⋅ Masoud Asgharian
Abstract
Autonomous language agents take irreversible actions, and many harms are compositional: a private read becomes a leak only once it reaches an external sink several actions later. Conformal risk control gives a finite-sample guarantee, but calibrating it against step-level harm forces a trade-off: loosen the threshold and compositional harm slips through, tighten it and benign work gets blocked. We introduce G-CARB (Graph-Localized Conformal Agent Risk Budget), which restricts the scorer to the source-to-sink dependency component of the proposed action while leaving the certified target global. On AgentDojo replay at $\alpha=.05$, a step-level rate target meets its own bound while episode harm runs $2.7$ and $3.1$ times over budget on two small open-weight backbones. Once the ledger is the calibrated unit, an operator who asks for $.05$ gets $.05$: G-CARB holds that budget on about half the scorer input, completes more tasks than full-prefix scoring at every operating point, and intervenes less wherever the budget leaves room to act.
Chat is not available.
Successful Page Load