Towards Self-Supervised, Generalizable and Decomposable 4D Driving Scene Reconstruction
Abstract
Reconstructing high-fidelity and controllable digital twins from multi-modal sensory observations is a fundamental problem in physical AI applications. While existing per-scene optimization methods handle dynamic driving scenes well, they are slow to optimize, have artifacts at novel viewpoints, and require costly manual annotations for decomposing the scene into controllable instances. Self-supervised generalizable reconstruction methods enable faster and more scalable reconstructions by learning from large datasets, but existing approaches do not fully leverage multi-modal inputs (e.g., camera and LiDAR) and lack decomposition capabilities. To address these limitations, we propose STRIDE, a self-supervised, generalizable and decomposable 4D driving scene reconstruction method. By processing multi-modal sensory inputs in a common 3D space with a Point-Transformer, STRIDE efficiently discovers spatio-temporal correspondences necessary for accurate flow prediction. By incorporating latent instance tokens, STRIDE is the first to perform feed-forward reconstruction and learnable decomposition without manual annotations. Experiments on public driving datasets show STRIDE achieves state-of-the-art performance on recovering geometry, appearance, and 3D flow. Moreover, we demonstrate how learned decompositions can enable dynamic instance manipulation and controllable simulation.