Modality-Autoregressive World-Action Models
Abstract
World-action models (WAMs) jointly predict future observations and actions, but typically represent the future using only RGB. Additional visual modalities such as depth, pretrained visual features, and point tracks provide complementary geometric, semantic, and motion information, yet effective multimodal WAM formulations are missing. We introduce ModAR, the first WAM to autoregressively generate multiple future modalities before generating actions. This allows each prediction to condition on previously generated modalities. Training from scratch, we contribute the first systematic study of training-data mixtures, future-observation prediction modalities, and WAM formulations. Our results suggest complementary gains from predicting point tracks, DINO features, and depth maps, while additionally predicting future RGB yields no consistent gain. ModAR convincingly outperforms prior WAM formulations, and achieves the highest average success rate at all evaluated data scales. On three real-world bimanual manipulation tasks, ModAR outperforms baselines and improves with both in-domain and out-of-domain human demonstrations, highlighting its ability to scale with cross-embodiment and out-of-domain data.