Red-Teaming Procedural Fairness in LLM Group Facilitators: Breaking the Process, Not the Model
Abstract
Large language models are increasingly used to facilitate group discussion, where their actions can affect who is heard, how contributions are represented, and how groups move toward consensus. These procedural functions make fairness of the group process an important concern. We study whether an ordinary participant can manipulate these functions in ways that undermine procedural fairness, without prompt injection or privileged access. To evaluate this threat, we define five attack scenarios covering participation, consensus, agenda setting, topic coverage, and attribution, and use multiobjective evolutionary search to generate plausible participant messages against GPT-4o and Claude Sonnet facilitators. Across 480 confirmation runs, standard facilitators exhibited targeted procedural breaches in nine of ten scenario--model combinations. Explicit procedural-fairness safeguards substantially reduced these failures but did not eliminate them under adaptive attack. Residual vulnerabilities differed across models and procedural properties. These findings identify procedural fairness as an adversarial attack surface for LLM facilitation and show that prompt-level safeguards alone do not guarantee the integrity of the group process.