MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
Abstract
Compositionality is critical for 3D object and scene generation, but existing part-aware 3D generation methods suffer from poor scalability due to quadratic global attention costs when increasing the number of components. In this work, we present MoCA, a compositional 3D generative model based on a novel component-level sparse attention that explicitly models inter-component dependencies. The attention mechanism features two key designs: 1) importance-based component routing that utilizes a lightweight router module for cross-component importance estimation and selects top-k relevant components for fine-grained interaction, and 2) distant components compression that preserves spatial priors while reducing computational complexity of global attention, by including compressed distant components into attention calculation instead of discarding them. With these designs, MoCA enables intricate compositional 3D asset creation with scalable number of components. Extensive experiments show MoCA outperforms baselines on both compositional object and scene generation tasks. Code and models will be made publicly available.