VLS: A Vision-Language-Shape Model for Open-Vocabulary Partonomic 3D Reconstruction
Abstract
Single-view 3D reconstruction has advanced rapidly in recent years, but existing research primarily focuses on recovering whole-object geometry while ignoring their semantic parts, which are crucial for fine-grained 3D perception and downstream applications. This paper introduces the task of open-vocabulary partonomic reconstruction. Given a single image and a set of part names, we aim to reconstruct both the object’s overall shape and its constituent semantic parts, even when the object and part categories are unseen during training. To address this task, we propose a vision-language-shape (VLS) model that unifies vision, language, and shape representations within a shared neural field and grounds continuous 3D coordinates to 2D pixel contexts via a deformable implicit function. It can effectively reconstruct any-topology shapes and open-vocabulary semantic parts from a single image. To train VLS with limited 3D part-labeled data, we propose an omni-supervised learning framework that leverages heterogeneous datasets with different levels of annotation and open-world knowledge from existing vision-language models. Extensive experiments on ShapeNetPart, PartNet, and Objaverse demonstrate the effectiveness and strong generalization ability of VLS. We will release the code and data publicly.