Learning World Models with Joint-Embedding Prediction for Single- and Multi-Agent Reinforcement Learning
Abstract
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal in both single- and multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations. A categorical latent state and causal Transformer are trained with posterior, action-conditioned dynamics, and masked spatial prediction objectives and are then used for actor-critic learning in latent imagination. For multi-agent control, we add a training-only joint predictor that conditions on synchronized local states and the joint action to predict each agent's next local observation representation. These predictions are passed through the same local posterior used during real interaction, while a centralized critic is used only for value learning; execution remains decentralized. On visual DeepMind Control, MA-JEPA is competitive with DreamerV3. On the SMAC benchmark, joint prediction improves performance over the matched independent local model while retaining decentralized execution. Our results show that joint-embedding prediction provides a practical learning objective for world models and extends to coupled multi-agent dynamics.