The Mask Is Not the Object: Volumetric Supervision for 3D Gaussian Segmentation
Abstract
Existing 3D Gaussian Splatting (3DGS) segmentation methods produce accurate rendered masks yet often fail to select a complete set of object Gaussians: removing the predicted Gaussians leaves the rendered depth in the target region nearly unchanged. We trace this gap to the contribution-weighted supervision path of rendering-based feature learning. Because Gaussian geometry and opacity are pre-optimized for RGB reconstruction, rendered-feature losses deliver strong gradients only to high-contribution Gaussians, leaving other Gaussians in the same object region weakly supervised and ultimately unselected. We propose VODA, a plug-and-play supervision objective that addresses this bias. After RGB optimization, Gaussians of the same object portion concentrate at similar depths along the camera ray. A depth-aware rasterizer exploits this property, grouping per-pixel contributing Gaussians into depth-coherent clusters and selecting the cluster that best explains each object pixel. The rendered feature of the selected cluster is then applied as a target directly to every Gaussian within it, bypassing the contribution-weighted gradient path. A complementary pixel-level loss further sharpens boundaries between adjacent objects. To evaluate volumetric completeness, we introduce Removal-3D and the Removal Depth Error (RDE), which measure whether removing the predicted Gaussians induces the expected depth change in 3D. Across four rendering-based baselines, VODA consistently improves volumetric segmentation on Removal-3D while maintaining or improving standard performance on LERF and ScanNet.