Turning Video Foundation Models into Action- Conditioned World Models by Steering Denoising
Abstract
Large pretrained video generators provide strong priors over appearance and motion, making them an appealing starting point for robotic world models. However, action conditioning them through finetuning is expensive and must be repeated for every new embodiment. Prior work showed that frozen video diffusion models can instead be action-conditioned by a lightweight adapter operating in prediction space, but only on smaller, outdated models. We find that directly transferring this recipe to a stronger frozen prior can improve video quality while leaving the action largely ignored. We identify three causes: learned mixture can suppress the adapter gradient, a competing full-prediction branch admits copying of the frozen prediction as a cheap solution, and action information is rewarded primarily at high noise. We introduce VIDeo GENeration STEerINg (VIDGENSTEIN), which replaces learned competition with an additive residual and concentrates adaptation where the action matters. On Wan2.2, VIDGENSTEIN substantially increases action dependence while retaining video quality and requires less training compute than an internal LoRA update. Experiments with DynamiCrafter and EasyAnimate further characterize how this prediction-space adaptation regime changes across diffusion and flow-matching priors.