Robot Learning with World Models: Capabilities, Frontiers, and Challenges
Abstract
Despite early successes of world models on robot learning, significant challenges persist in physical accuracy, spatiotemporal consistency, and the lack of critical modalities beyond vision (e.g., tactile sensing, proprioception). This workshop aims to bring together researchers from generative modeling, computer vision, and robotics to exchange ideas on (1) the state-of-the art of robot learning with world models (where are we now?), and (2) the open challenges and gaps (where should we go?). Through invited talks, panel, oral presentations and poster sessions, we will discuss topics including world action models, improving physical accuracy, non-vision modalities, evaluation metics, and learning and planning in imagination. The workshop welcomes contributions ranging from short papers, full papers, demos, and networking group proposals. We hope to create an inclusive venue for active idea exchanging and community building among people with diverse backgrounds.