StaDy: Factorizing the World into Static and Dynamic via Likelihood Matching
Abstract
Latent action models (LAMs) aim to learn compact representations of state transitions directly from visual observations, but they suffer from a fundamental ambiguity: the latent code can encode target-state information rather than true dynamics. Existing approaches address this issue through restrictive bottlenecks, which reduce leakage at the cost of limiting expressivity. We propose StaDy, a regularization framework that instead enforces conditional informativeness: the latent code should aid prediction only when paired with the source state, while remaining uninformative about the target on its own. Concretely, we introduce a likelihood-matching objective that aligns the decoder’s predictions conditioned solely on the latent variable with its unconditional predictions. This discourages target memorization without constraining latent capacity. Experiments in video modeling show that, unlike bottleneck-based methods, StaDy scales effectively with increased latent dimensionality, reducing static content leakage while improving the representation of more complex dynamics.