SVG-3D: Mining Decision Boundaries with Generative Splatting Priors for Zero-Shot 3D Classification
Ryo Umagami ⋅ Shohei Ohsawa
Abstract
Zero-shot 3D point cloud classifiers built on large-scale pretrained models (OpenShape, Uni3D, ReCon++, DuoduoCLIP) excel on clean synthetic benchmarks such as ModelNet40 but degrade sharply on real-world scans (ScanObjectNN, ScanNet) and corrupted point clouds (ModelNet-C)–a persistent real-world generalization gap. We propose SVG-3D, the first 3D extension of Support Vector Generation (SVG), to close this gap without retraining a more capable encoder or accessing target-domain test data. SVG-3D mines a discriminative kernel-machine support set from a frozen encoder paired with DiffSplat, a text-conditioned 3D Gaussian splatting generator. It rests on two design choices that preserve strict zero-shot rigour: (i) a dynamic pair-selection rule that picks the top-$K$ most confusable class pairs from a pseudo-confusion matrix computed entirely on generated samples, never on the test set; and (ii) a Metropolis–Hastings sampler operating on the CLIP text-embedding hypersphere with a Slerp-plus-Gaussian proposal whose density admits a closed-form ratio, yielding the same acceptance rule as SVG. On ScanObjectNN (OBJ\_ONLY/OBJ\_BG/PB\_T50\_RS) and ScanNet, SVG-3D surpasses prior zero-shot 3D classifiers built on large-scale pretrained models without consulting any test data; on ModelNet-C, a complementary corruption-aware setting in which the corruption family is assumed known further improves mean robustness over the same baselines.
Chat is not available.
Successful Page Load