Pygmalion: Bridging Reconstruction and Generation in Sparse Voxel-based 3D Modeling
Abstract
Recent advances in sparse voxel-based VAEs have demonstrated remarkable capabilities in high-fidelity 3D autoencoding, yet their generative counterparts consistently lag behind their reconstruction performance. A fundamental cause of this gap lies in the \emph{brittle output representation} in sparse voxel decoders, which parameterize geometry through discrete topological decisions, such as intersection flags or occupancy signs. Small perturbations in these parameters, which are unavoidable under diffusion sampling, can be amplified into abrupt topological ruptures and grid-like artifacts. We argue that a generation-friendly representation should ensure that such perturbations induce only smooth geometric transitions. To this end, we propose Pygmalion, a generation-friendly autoencoding framework that reformulates sparse voxel decoding as a coupled SDF parameterization. To retain explicit mesh-level supervision within this parameterization, which is non-trivial as the SDF jointly governs both topology and geometry, we introduce hinge-based sign correction, case-aware geometry supervision, and rendering-based refinement, achieving high-fidelity reconstruction while preserving generative robustness. Built upon this generation-friendly representation, Pygmalion scales consistently across DiT model sizes and voxel resolutions. Extensive experiments demonstrate state-of-the-art reconstruction fidelity and high-quality 3D generation with complete surfaces, sharp features, and fine geometric details.