PhysFlow: Physics-Intrinsic Velocity Regularization for Motion-Intensive Video Generation
Xianglong Guo ⋅ Chang Yu ⋅ Haobo Xu ⋅ Junhao Ma ⋅ Zhen Lei
Abstract
Text-to-video models built on flow matching generate compelling videos for common prompts but degrade systematically under motion-intensive scenarios, producing artifacts such as object fragmentation, geometric deformation, and temporal flickering. Existing training-free approaches operate at the attention or guidance-scale level, neither of which directly addresses the velocity field that governs latent motion evolution. We observe that these motion artifacts stem from local misalignment between the predicted velocity and the latent's frame-axis temporal structure, a signal that can be diagnosed from quantities the sampler already computes. Based on this, we propose PhysFlow, a training-free velocity regularization method. PhysFlow constructs a per-frame latent flow residual to locate motion-inconsistent regions, steers the velocity to restore frame-axis alignment, and bounds the correction via residual-aware masking and safe clipping. We also introduce MotionStress-100, a five-category benchmark with a VLM-based protocol that isolates motion-intensive failures. On Wan2.1, PhysFlow improves the average MotionStress score over the baseline with no additional inference cost ($1.00\times$ cost), outperforming CFG-Zero* (-10.0%, $0.98\times$ cost) and FlowMo (+0.7%, $2.12\times$ cost) by at least 4.7%, while improving VBench motion smoothness without loss of visual fidelity.
Chat is not available.
Successful Page Load