SmolWAM: A Latent World-Action Model for Robotic Manipulation on Edge Devices
Oluwatimilehin Owolabi ⋅ Hikmah Olawore ⋅ Frances Adelakun ⋅ Ayotomiwa Oyewumi
Abstract
World--action models (WAMs) are robot policies built on top of video-generation models, and they work very well: NVIDIA Cosmos Policy reaches 98.5\% success on the LIBERO manipulation benchmark. They are also expensive to run. The released 2B-parameter model needs 1.8 seconds to produce 16 robot actions on an NVIDIA L4 GPU, which is too slow for real-time control and too heavy for the computers robots actually carry. The usual fix is to train a smaller policy from scratch, but that gives up the video prior and costs success. We take a different route: we start from the released checkpoint and remove the parts of inference the robot does not actually need. The result is \textbf{SmolWAM}, a latent world--action model. It keeps the teacher's training recipe but acts entirely in latent space: one denoising step instead of five, future frames left as noise instead of being predicted, camera images encoded as single frames instead of 33-frame clips, and no decoding of predicted video. These changes alone make the unchanged 2B model $8\times$ faster with no measured loss in success. We then remove 8 of its 28 transformer blocks and repair the smaller model with about one hour of LoRA distillation on 414 observations from the teacher's own rollouts. The resulting SmolWAM-1.4B reaches 87.5\% on a 400-episode LIBERO protocol at 78 actions per second, with $7.3\times$ fewer FLOPs and 3.1\,GiB of memory, small enough for embedded GPUs; every experiment ran on a single 23\,GB L4. A WAM's imagined future is needed for training, not for acting: most of the cost of video generation can be dropped at inference with no measured loss in success.
Chat is not available.
Successful Page Load