Belief-conditioned Predictive Latent Embedding for Zero-shot Multi-agent Coordination
Abstract
Zero-shot coordination (ZSC) requires an agent to adapt to partners whose behavior was not seen during training. Most existing methods improve ZSC by increasing partner diversity, but this leaves a representation problem unresolved: from the ego-agent's perspective, the same state can imply different futures depending on the unobserved partner policy. When rewards are sparse or observations are high-dimensional, reward-driven training provides little signal for preserving these partner-dependent futures. In this work, we propose belief-conditioned successor features (BSFs), an SF formulation for decentralized multi-agent settings with latent partner types. By conditioning on sufficient information about beliefs inferred from interaction history, BSFs lift the ego-agent's process to a Markov state space and define successor features over partner-dependent future occupancy. We implement this idea in MA-JEPA, an on-policy self-play (SP) algorithm that uses a learned history encoder and an auxiliary latent prediction loss to predict discounted future latents. Across Overcooked settings covering shaped rewards, sparse rewards, RGB observations, as well as SMAC N-agent ZSC, MA-JEPA outperforms SP, population-based, and agent-modeling baselines. Specifically, it achieves performance gains in sparse rewards and RGB observations setups by about +77% and +32%.