MilliVid: Adaptive Latents for Long-Range Consistency in Video Generation
Abstract
Transformer-based video generative models have become increasingly realistic, but long-horizon consistency remains challenging to achieve because even a few dozen frames create impractically long sequence lengths. We show that this issue can be mitigated by generating video using coarse-to-fine rollout within a multi-scale token space. Our approach is simple: first, we pre-train an adaptive autoencoder that compresses each frame into a hierarchy of tokens, with levels ranging from the typical latent resolution to only a handful of tokens per frame. This yields a hierarchy in which the coarsest levels capture the most consequential information—such as scene layout and semantics—while finer levels add high-frequency appearance and texture. Then, we train a video diffusion model to generate these tokens using coarse-to-fine rollout. By carefully controlling the level of detail at which frames are generated and used as context during each rollout step, we are able to preserve long-range consistency in geometry and object permanence while expending less compute on long-term consistency of less perceptually relevant details. We validate this approach using a custom long-horizon Minecraft dataset, where it produces substantially more consistent rollouts compared to strong baselines. We hope this perspective opens new opportunities for long-term consistent video generation.