Cooperative agents specialize more but explore less
Abstract
Cooperation helps multi-agent systems solve complex tasks, but its benefits may not last. We find that cooperative agents specialize more but explore less, which eventually hurts group performance. We study this process with deep reinforcement learning agents and large language models under various contexts of sequential social dilemmas. In all settings, agents must choose between producing rewards for themselves and maintaining resources to help their group. Compared with individual incentives, collective incentives lead agents to cooperate more and perform better. Cooperation goals also cause initially homogeneous agents to take on maintenance and production roles, even though we never assign these roles. Once agents specialize, maintenance agents explore significantly less than production agents and become progressively less exploratory. As maintenance agents narrow their search, they fall short of providing sufficient resources. Production agents then have fewer resources to use, and group performance can decline by up to 90\%, long after agents learn to cooperate. Robust multi-agent systems must therefore not only learn to cooperate, but also keep exploring after cooperation succeeds.