PlasticMem: Adding Temporal Reasoning to Diffusions for Consistent Long Video Generation
Abstract
Scaling video generation from short to long durations has attracted growing attention as a way to democratize video creation. The central challenge lies in long-term consistency, which requires stable rollout over extended horizons and accurate visual memory of previously generated elements. However, the scarcity of high-quality long video data remains a fundamental bottleneck, forcing prior methods to rely on suboptimal synthetic data that either limits dynamic range and visual memory or suffers from a large simulation-to-reality gap. To address this, we explore a new generation paradigm, Temporal Reasoning-based Rendering (TRR), which better exploits existing short video data for consistent long video generation. Under TRR, we propose PlasticMem, which plans contents from scripts, then reasons about and independently renders each action to form a consistent long video. Benchmark results show that PlasticMem, trained on short videos, achieves stable minute-scale long-horizon generation and visual memory.