Gated Transformer World Models for Model-Based Reinforcement Learning
Abstract
Model-based reinforcement learning has emerged as a promising approach for training sample-efficient policies by learning compact latent dynamics models that simulate, plan, and act through imagination. Recent work has shown that transformer-based world models outperform recurrent architectures due to their ability to capture long-range temporal dependencies. However, standard self-attention mechanisms apply the attention score uniformly across all dimensions of the token's representation, lacking an explicit mechanism to regulate information flow across dimensions. We introduce a gating mechanism for transformer-based world models that implements a feature-wise modulation mechanism within the attention computation. The gating module adaptively scales dimensions in a data-dependent way, enabling the model to selectively control contextual information. Our gating mechanism achieves greater sample-efficient policy learning through normalized mean performance improvements for the Atari 100k benchmark when applied to the STORM and TWISTER transformer world models.