Helpful Allies, Better Liars: RL-Trained Agents Meet Human Social Deduction Experts
Yannik Keller ⋅ Levin Brinkmann ⋅ Thomas F Eisenmann ⋅ Iyad Rahwan
Abstract
AI agents are increasingly deployed in shared environments where they must infer the goals of humans and other agents to identify allies, opponents, and attempts to mislead them. Existing work has shown that post-trained instruct LLMs often exhibit misaligned behavior such as deception in these settings. However, it remains unclear how social and anti-social behaviors emerge across training stages. Here, we address this gap by finetuning a pretrained base model to play the hidden-role linguistic social deduction game ``Secret HAL'' through imitation and reinforcement learning (RL). We evaluate both checkpoints against 34 highly experienced players (mean $\sim$3,484 prior games) in 192 real-time games with mixed human-AI compositions, in which turn-taking is unconstrained and no player is told which seats are AI. We find that imitation-only agents trained on 21,459 human game records lack the strategic and deductive skills to be helpful to human expert teams and utter fewer lies than human experts. However, RL optimization for outcome yields stronger strategic skills, showing deductive capabilities comparable to those of human experts. On the flip side, the rate at which these agents make false claims in the traitor role rises to human expert level, despite deception never being directly rewarded. More broadly, this result reveals a dual effect about RL-based LLM post-training. RL on outcome is useful for producing helpful cooperators, but can also amplify deceptive behavior.
Chat is not available.
Successful Page Load