Youdunit: Single-Call Counterfactual Necessity in Multi-Agent LLM Systems
Marissa Li ⋅ Stephanie Gao ⋅ Kenny Guo ⋅ Xingjian Li ⋅ William Chang ⋅ Goran Radanovic
Abstract
Agentic systems built on large language models (LLMs) are increasingly deployed in high-stakes settings, but the causal mechanisms by which adversarial agents induce harmful outcomes remain poorly understood. We address this with a single-call counterfactual necessity test for multi-agent LLM systems that asks: \emph{which specific agent communication was causally responsible for the harmful outcome?} In contrast with actual-causality frameworks that reason over set-valued causes, we use one attribution rule throughout: single-call necessity for a selected proximate call. The key technical contribution is the \emph{Gumbel-Max tape}, which lifts the per-token Gumbel-Max counterfactual generator of \cite{chatzi2024counterfactual} from a single LLM to a multi-agent conversation. Per-token GPU random number generator (RNG) states are recorded during a factual run on a tape shared across all agents, then replayed during counterfactual (CF) runs everywhere except at the intervened call. Applied to the BAD-ACTS benchmark across $148$ adversarial scenarios spanning four multi-agent environments and two prompting conditions (\emph{non-safe} and \emph{safe}), the test shows that single-call necessity tracks communication structure: decentralized and hierarchical environments (Travel Planning, Financial Article Writing) exhibit meaningful aggregate causal effect (ACE), while sequential debate (Multi-Agent Debate) shows near-zero ACE despite comparable attack success rate. A replay sanity check confirms that all CF effects under identity intervention are exactly zero, validating the tape mechanism. The results identify where targeted defenses are likely to be effective in deployed agentic systems.
Chat is not available.
Successful Page Load