Scenes as Objects, Not Primitives : Instance-Structured 3D Tokenization from Unposed Views
Abstract
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be recovered after the fact. We propose a feed-forward framework that directly decomposes unposed multi-view images into instance-structured 3D token groups - compact object-centric units from which reconstruction, segmentation, and manipulation all follow. Each token group consists of an instance token that captures entity-level identity, paired with anchor tokens that model local geometry and appearance and decode into 3D Gaussians. This factorization separates what belongs together from how each part looks, making object instances a native interface of the representation rather than a derived product. The model is trained through differentiable rendering with joint reconstruction and segmentation supervision, without 3D annotations or test-time optimization. On indoor scene benchmarks, our feed-forward model outperforms per-scene optimized baselines in class-agnostic instance segmentation, while the same token groups support instance-level scene editing through direct token removal, translation, and insertion.