Do Deep Stereo Models Learn Shape Priors?
Abstract
Localizing object boundaries is important for embodied visual systems, since small boundary errors can change the inferred geometry of physical surfaces. We study whether deep stereo models learn shape priors that allow them to recover precise boundaries when local image and disparity cues are weak. We create a synthetic stereo dataset consisting of objects drawn from a highly constrained family of shapes, with randomized depth, slant, color and texture. Despite the simplicity of this shape space, we find that several state-of-the-art stereo models fail to fully exploit its structure, even when finetuned or trained from scratch on the dataset. We then introduce a simple generative stereo model that jointly predicts disparity and a distance field representing object boundaries. The model improves boundary localization and produces contours that adhere more closely to the underlying shape space. These results expose limitations in the shape priors learned by current stereo architectures and suggest that generative models are a promising means of combining regional grouping evidence with global geometric coherence.