Spatially-Grounded Long Video Generation with Self Geometry Forcing
Abstract
Recent video generative models can produce photorealistic short clips, but struggle with long-horizon generation under dynamic camera control, often suffering from viewpoint drift, geometric inconsistency, and error accumulation. We propose Joint Appearance-Geometry Diffusion Transformer (JAG), a spatially-grounded causal video diffusion model that explicitly integrates 3D perception for long and consistent video generation. It adopts a novel dual-branch causal architecture that jointly models visual appearance and scene geometry. Building upon this formulation, we introduce Self Geometry Forcing, a geometry-aware distillation scheme that conditions current generation on self-predicted geometric structures as persistent long-term memory and imposes self-verified geometric rewards on intermediate rollouts, effectively mitigating error accumulation over long horizons. Extensive experiments show that JAG can generate long videos with superior coherent appearance and geometry at constant per-frame computation, improving long-horizon controllability, geometric consistency, and visual fidelity over previous state-of-the-art methods.