Can LLMs Deliberate? Benchmarking Collective Reasoning on Normatively Contested Problems
Abstract
Existing benchmarks for multi-agent LLM systems evaluate collective reasoning on tasks with verifiable solutions, such as mathematics or coordination games. However, LLM agents are increasingly proposed for democratic applications, where collective reasoning is normatively contested and solutions cannot be assessed objectively for correctness. To benchmark LLMs' deliberative capabilities in such contexts, we introduce DelibSim, a configurable simulation environment that evaluates multi-agent deliberation along three theoretically grounded dimensions adapted from political science: procedural discourse quality (AQuA), deliberative reasoning quality (DRI), and perspective diversity as a key factor for the epistemic value of deliberation. DelibSim spans 12 real-world citizen-assembly topics with matched human reference data, three prompting regimes, and 11 frontier model configurations, including reasoning-enabled and mixed-model ensembles. Across 1,980 five-agent deliberations, we document a gap between discourse and reasoning quality. LLM groups achieve discourse quality statistically indistinguishable from human deliberation (AQuA 2.94 for LLMs vs. 2.98 for humans), while normative prompting yields a small but significant gain in shared understanding (delta-DRI = 0.029 for LLMs vs. 0.099 for humans). Yet the results show substantial topic heterogeneity and turn negative on ethically complex topics. Most strikingly, LLM groups exhibit far lower perspective diversity than human groups (6.5 vs. 18.8) and reversed convergence dynamics: human deliberation decreases dispersion as diverse views synthesize, whereas LLM deliberation increases it. DelibSim thus exposes a failure mode invisible to standard evaluation: surface-level discourse quality can mask fundamentally different reasoning dynamics. We release DelibSim as an open benchmark together with the generated simulation results.