Scalable Multi-Agent Contrastive Reinforcement Learning
Abstract
Sparse rewards make cooperative multi-agent reinforcement learning difficult because useful feedback often appears only after multiple agents coordinate over long horizons. We introduce Multi-Agent Contrastive Reinforcement Learning (MACRL), a goal-conditioned method that trains parameter-shared decentralized actors with a centralized contrastive critic. Instead of learning from sparse task rewards, the critic contrasts joint state-action embeddings with future achieved goals from replay, turning every trajectory into dense self-supervised signal. Across tasks that vary in horizon length, coordination difficulty, and object interaction, our method is competitive with strong HER-based baselines when exploration is sufficient and substantially more robust when sparse rewards become uninformative. On the hardest tasks, it obtains substantial non-zero success while non-contrastive baselines remain near zero. Scaling from 4 to 64 layers further improves performance on several hard tasks, with gains up to 25 percentage points, while providing little consistent benefit to MA-SAC+HER. These results suggest that contrastive objectives and depth-scaled residual networks provide a promising foundation for scalable cooperative goal-conditioned control.