Who, Where, and What? Forensic Localization in LLM-Based Multi-Agent Systems
Abstract
LLM-based multi-agent systems extend single agents with role specialization and inter-agent communication, but they also expand the attack surface: an attacker can compromise an externally writable module of one agent, and the malicious payload may propagate across communication hops before some downstream agent invokes an attacker-chosen tool. Existing defenses are routinely circumvented by stronger or adaptive attacks, which motivates a complementary forensic capability that traces an incorrect tool invocation back to its root cause. We formulate this problem as three-level forensic localization, where the goal is to jointly identify the source agent whose internal state was first compromised, the specific module of that agent that was exploited, and the exact units inside that module carrying the malicious payload. To enable systematic evaluation, we construct MAFL-Bench, the first forensic localization benchmark for LLM-based multi-agent systems, which spans three application scenarios, four communication structures, and twelve attack instances, with 21,200 interaction logs annotated at the agent, module, and unit levels. We further propose MATracer, a black-box forensic framework that resolves the three-level attribution through a coarse-to-fine cascade driven by log-likelihood queries on a proxy LLM, progressively localizing the source agent, the compromised module, and the contaminated units. Extensive experiments show that MATracer accurately localizes the attack across diverse multi-agent settings, consistently outperforms ten baselines, and remains robust under adaptive attacks.