Safe Actions Can Form Unsafe Traces: Benchmarking and Shielding Compositional Emergent Risk in AI Agents
Abstract
Large language model agents increasingly act through tools, browsers, code interpreters, and external APIs, turning safety from a single-output problem into a trace-level problem. We identify \emph{compositional emergent risk} (CER), a failure mode where individually safe actions interact across time to produce unsafe outcomes. We show that bounded-window safety filters can miss CER whenever the risky dependency lies outside their visible context. To study this failure systematically, we introduce \textsc{CER-Bench}, a controlled long-horizon benchmark with 440 tasks and 19,525 actions across 5--500 steps, spanning five risk domains and two compositional mechanisms. At the largest tier, each trace induces a 79M+ candidate risk-composition search space. Across 15 frontier and open-source LLM agents, every model exhibits a non-zero compositionality gap (CG), ranging from 44.4\% to 100.0\% with a mean of 82.2\%. Moreover, 86.1\% of compliant executions contain caution language yet still proceed, showing that verbal risk awareness does not reliably prevent unsafe composition. We introduce \textsc{RiskShield}, a conformal-calibrated runtime shield that learns trace-risk boundaries, evaluates planned continuations before commitment, and substitutes risky suffixes while preserving safe prefixes. On 7 held-out \textsc{CER-Bench} models, \textsc{RiskShield} reduces mean \textsc{CER-5} CG from 69.9\% to 0.0\%, with no shielded pass observed in 763 evaluations and a 95\% upper confidence bound of 1.6\%, while preserving 78.4\% step-level task completion. It outperforms cumulative-threshold, sliding-window, full-trace, and published safety baselines, with 63.3\% of tasks requiring no extra API call and additional robustness across longer traces, risk domains, and external agent-safety benchmarks. Code and benchmark are available at \url{https://anonymous.4open.science/r/RiskShield-2284}.