NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games
Tomáš Holeček ⋅ Viliam Lisy
Abstract
Model-based reinforcement learning (MBRL) has achieved remarkable sample efficiency in single-agent domains, yet its extension to competitive imperfect information games (IIGs) remains underexplored. In multi-agent settings, opponent-induced non-stationarity complicates the learning process, and decentralized model learning faces severe identifiability barriers. To solve this, we propose NashDreamer, a principled MBRL framework for two-player zero-sum IIGs. NashDreamer introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that decouples environment dynamics from the effect of players' strategies on their individual observations. For policy optimization, NashDreamer uses Regularized Nash Dynamics (RNaD) within the latent imagination, providing theoretical convergence guarantees toward Nash equilibria. Empirical evaluations across Imperfect Information Goofspiel, Leduc Hold'em and Battleship demonstrate that NashDreamer achieves vastly superior sample efficiency compared to model-free baselines, including a $71.0$\% head-to-head win rate against model-free RNaD in stochastic 13-card Imperfect Information Goofspiel. Finally, we theoretically analyze the architecture's optimization landscape, formally identifying the vulnerability of the Dreamer family of algorithms to posterior collapse in highly stochastic environments, and we highlight the open challenge of mitigating latent non-stationarity without enabling degenerate solutions.
Chat is not available.
Successful Page Load