SciFusion: Compositional Representations for Physics-Guided 3D Scientific Image Generation
Abstract
Existing 3D generative models for scientific imaging are typically trained separately for each modality, limiting knowledge transfer across heterogeneous imaging domains and requiring substantial modality-specific data. We present SciFusion, a physics-guided compositional latent diffusion framework that learns a unified representation for 3D scientific image generation across modalities. SciFusion factorizes each modality into reusable domain, object, and contrast representations, allowing shared structure to transfer across MRI, CT, and electron microscopy (EM) while preserving modality-specific characteristics, and enabling zero-shot generation of unseen modality compositions without target-modality training. To respect modality-dependent image formation, we incorporate domain-specific physics guidance based on MRI power-spectrum statistics, CT Radon projections, and EM diffraction patterns. Experiments across 15 medical and scientific imaging modalities show that SciFusion improves generation quality under unified training and successfully synthesizes unseen modality compositions. These results suggest that compositional, physics-grounded representations provide a scalable framework for transferable 3D scientific image generation.