Cooperative Words, Selfish Actions: How Strategic Optimization Shapes LLMs' Interaction With Humans
Abstract
Large language models (LLMs) are increasingly optimized to act and communicate in social environments, but how do their objectives reshape the strategies they use with humans? We study this question in four repeated economic games with natural-language communication. Starting from an LLM trained to imitate human actions and messages, we use reinforcement learning to optimize agents either for their own payoff or for a prosocial objective that also penalizes payoff inequality. Both objectives increase agents' rewards, yet produce opposite consequences for their human partners. Selfish agents outperform humans while reducing human payoff and sacrificing collective efficiency and fairness; prosocial agents instead achieve high rewards through more efficient and equitable interactions. Crucially, these objectives also reshape communication. Both RL-trained agents use more cooperative, relational, joint-goal language, and their messages can increase human cooperation. But similar language masks sharply different strategies: selfish agents become substantially more likely to state intentions that contradict their subsequent actions, despite receiving no explicit reward for deception, whereas prosocial agents use communication to establish welfare-improving coordination and are as honest as or more honest than humans. Linear probes further show that strategic optimization strengthens representations of the co-player's future return, most strongly in selfish agents, separating social capability from prosocial objective. Our results show that socially cooperative language need not imply cooperative incentives and can instead mask behavior that is misaligned with human welfare. Optimizing social agents therefore shapes not only what they do and how they communicate, but how they influence the humans with whom they interact.