When the Most Disruptive Messages Are Safe to Prune in Aggregate: Functional Analysis of Communication Pruning in Two-Agent LLM Systems
Abstract
Multi-agent LLM systems exchange messages whose communicative necessity varies widely, yet no systematic evaluation protocol determines which messages can be safely removed. We introduce, to our knowledge, the first message-level functional analysis protocol for evaluating communication in two-agent LLM systems. The protocol comprises three reusable stages (annotate, perturb, prune), each deployable to new backbones, domains, and topologies without task-specific training data. We release a 21,661-message annotated corpus across 3 backbones and 4 reasoning domains with a five-class taxonomy for mechanism exploration, a zero-shot annotation pipeline using constrained decoding, and a statistical hypothesis-testing framework combining TOST equivalence, non-inferiority, and McNemar's tests. Human validation (N=225) confirms annotation reliability (κ = 0.81 binary ack/non-ack, underpinning all pruning claims; 0.76 five-class); the cross-model annotation pipeline achieves κ=0.72. Applying the protocol across 18 configurations spanning 7B–72B scale, 4 reasoning domains, and 3 topology variants, with 10-seed replication for flagship settings, we surface a perturbation–pruning dissociation: the most disruptive message type under single removal is safe to remove in bulk, enabling acknowledgment-targeted pruning that suppresses 36–56% of inter-agent messages from conversation context at non-inferior accuracy under greedy decoding. The protocol further identifies boundary conditions where pruning fails and reveals backbone-dependent communication mechanisms, providing a screening tool for deployment decisions.