Decodable, Not Transferable: Latent Actions Across Worlds
Abstract
Latent action models aim to make unlabeled video useful for learning controllable world models by inferring latent actions from visual transitions. For these latents to serve as reusable action representations, however, their semantics must remain consistent across changes in visual context and environment. We ask what action-related structure can be identified in pretrained latent action model representations, and whether this structure can support action prediction in known and unknown environments. We answer it as three nested questions: (i) how is action information \emph{encoded} in the latents, (ii) is it \emph{decodable} into actions inside a seen environment, and (iii) does the same decoder \emph{transfer} to an environment never seen during probe fitting? We evaluate two recent latent action models, AdaWorld and Olaf-World, on a 193-game Stable Retro corpus and a subset of Open Pixels2Play. We find that action-label information is present, but it is weakly linearly separable and entangled with game-related variation. There is a substantial drop in action prediction performance on unseen environments compared to known environments. Moreover, a generic video tokenizer carries the strongest \emph{within-game} action signal of any representation we test, suggesting that such decodability may reflect generic visual dynamics rather than an action-specific inductive bias of latent action models.