SoDeArena: A Socially-Situated Reasoning Benchmark for Large Language Models
Abstract
In recent years, Large Language Models (LLMs) have demonstrated remarkable performance in tasks such as mathematical reasoning and code generation. However, as LLMs are increasingly deployed as autonomous agents, their reasoning capabilities in dynamic, human-like social interactions remain largely unexplored. To bridge this gap, we propose the benchmark SoDeArena—Social Deduction Arena, designed to evaluate the socially-situated reasoning capabilities of LLMs through a diverse set of interactive games. Through large-scale multi-agent simulations and human interaction experiments, we have identified three key insights for improving current LLMs: (1) the overall proficiency of models in social deduction is fundamentally bottlenecked by model scale and reasoning capabilities; (2) models lack proactive intent management, defaulting instead to strategic passivity and output conformity during multi-agent interactions; and (3) LLMs exhibit dual behavioral biases, manifesting as a rigid safety alignment in adversarial roles and a pathological reliance on shallow positional patterns over robust reasoning processes. These findings point to potential directions for the future development of AI systems with robust socially-situated reasoning capabilities.