Splatting the Invisible: Geometry and Appearance Scene Completion from Sparse Views
Abstract
Novel view synthesis methods such as 3D Gaussian Splatting degrade under sparse-view settings, suffering from an inability to extrapolate beyond observed regions. To enable high fidelity reconstruction and rendering from sparse views, inpainting of sparsely observed areas and completion of occluded regions is required. Unlike existing approaches that perform these tasks using 2D inpainting diffusion models or geometric-only 3D completion models, we propose a fully 3D scene completion framework, called InvisiSplat, that performs geometry and appearance completion directly in 3D. Our method consists of a Gaussian Splat VAE which learns a sparse point cloud latent space that predicts Gaussian Splat parameters from the latents, as well as a two-step flow model that learns to generate samples conditioned on the point cloud latents, different from other approaches that use tokens without 3D coordinates. Our flow model utilizes alternating local 3D attention and global attention, which balances the needs for precise local geometry and large-scale spatial context. Across several datasets, we demonstrate that our method produces 33% - 82% higher fidelity 3D geometry over prior work while attaining competitive performance on appearance metrics when compared to existing approaches.