Geometric Confidence as an Uncertainty Oracle for World-Model Perception: When Does Monocular Depth Help Classical Stereo?
Vidhi Kulkarni ⋅ Hang Zhao ⋅ Tejas S Anand ⋅ Victor Murta ⋅ Jing Du
Abstract
Robot world models depend on reliable perception, yet they rarely quantify *when* their perceptual inputs can be trusted, a prerequisite for detecting uncertainty in imagined rollouts before deployment. We study whether the calibrated geometric confidence produced by classical stereo matching can serve as an uncertainty oracle for hybrid perception, gating between deterministic classical depth and flexible monocular neural depth. Through a controlled empirical study on KITTI Stereo 2015, we establish three findings. First, classical stereo confidence is well-calibrated: pixels it flags as reliable exhibit $48%$ lower depth error than those it flags as uncertain. Second, despite this calibration, naive confidence-gated fusion fails to improve over a classical-only baseline; accuracy degrades monotonically as monocular weight increases. Third, scaling the monocular model $13\times$ ($24.8$M$\rightarrow$$335$M parameters) reduces low-confidence error by only $1.1%$, leaving classical stereo $30%$ more accurate even within its own low-confidence regions. We formalize the precondition any auxiliary signal must satisfy for fusion to help, and argue that calibrated geometric confidence is a deployable, training-free uncertainty signal for world-model perception. Our results caution that confidence calibration alone does not guarantee useful fusion, a diagnostic practitioners should apply before integrating auxiliary perception into world models.
Chat is not available.
Successful Page Load