StereoSplat: Metric-Scale Novel View Synthesis via Stereo-Grounded Gaussian Splatting
Abstract
We present StereoSplat, a feed-forward 3D Gaussian Splatting architecture designed explicitly for stereo videos to achieve reliable, metric-scale novel-view synthesis. Unlike monocular approaches that neglect the nature of binocular pairs, StereoSplat adopts a native stereo architecture that extracts dense multi-scale latent features and disparity from a foundation stereo backbone and feeds them into a dedicated Gaussian decoder. Our decoder utilizes cross-view feature fusion and gated spatial refinement to directly translate these rich representations into per-pixel 3D Gaussians. To resolve the high redundancy of per-pixel predictions across overlapping views, we propose a learnable depth-adaptive aggregation mechanism that clusters and aggregates primitives in inverse-depth space while preserving fine-grained details. To support robust training and generalization, we introduce XRStereo, a large-scale synthetic dataset of stereo video featuring various rig configurations tailored to match the physical hardware of real-world platforms. Trained exclusively on synthetic data and evaluated across challenging benchmarks, StereoSplat yields substantial improvements in the depth accuracy of rendered novel views over established baselines while maintaining high-fidelity novel view synthesis. By overcoming prior geometric limitations, our framework enables robust zero-shot transfer to real-world domains, offering a scalable geometric foundation for spatial computing in mixed reality, autonomous driving, and robotics.