Unifying Partner and Environment Diversity to Improve Human-AI Coordination
Abstract
Real-world applications of modern AI systems require generalization to both unseen environments and novel human users. The field of zero-shot coordination (ZSC) focused primarily on deploying AI to novel partners in otherwise familiar environments seen in the training. The popular population-based methods enable partner generalization by training the agent on a diversity of simulated human users (serving as training-time partners) in a fixed environment, aiming to cover the diverse behaviors of novel human users. Such a paradigm does not consider the more challenging realistic scenarios where both the users and the environment are unseen during training. Recently, training across environments has been explored in training a single self-play agent, which shows better generalization on both novel human users and unseen environments. In this work, we ask whether jointly modeling partner diversity and environment diversity yields stronger coordination with novel partners in novel environments. We introduce Cross Environment Co-Play (CECP), a two-stage framework that first trains a population of self-play agents across a distribution of procedurally generated environments, and then trains a cooperator policy against this partner population. To further couple social and environmental reasoning, we add a next-state prediction module that encourages representations to encode both partner behavior and task structure, and validate the design choice with ablation studies. We evaluate CECP on both a neuroscience-inspired cooperative foraging task, and the population ZSC benchmark Overcooked, under held-out partner and held-out environment conditions. CECP achieves the strongest performance among the compared baselines and attains state-of-the-art results under our evaluation protocol for human-AI collaboration. We provide further analysis using the model's latent representation encoded by the next-state prediction module, showing it can not only improve the performance of the cooperative AI, but also provide more interpretable latents, which demonstrate the inference of partner skill levels.