SOAR: Semantic Organ-Aware Pretraining for 3D CT Image Understanding
Abstract
Vision-language pretraining has advanced multimodal medical AI, but its direct application to 3D computed tomography (CT) remains limited by a fundamental mismatch between dense volumetric anatomy and sparse diagnostic reports. Existing 3D CT vision-language methods typically rely on global image-report alignment or align anatomical regions with decomposed report descriptions. However, a CT scan captures rich, spatially distributed information across many organs, whereas a report summarizes only a small subset of clinically relevant findings. As a result, report-based alignment provides incomplete or weak supervision, leaving many local anatomical structures underrepresented. Motivated by this limitation, we introduce SOAR, a semantic organ-aware pretraining framework for 3D CT image understanding. The key innovation of SOAR lies in its integration of fine-grained, organ-level structural priors, enabling report-free organ-aware supervision for visual pretraining without relying on radiologist annotations or coarse global image-report matching. SOAR combines three complementary objectives: (i) organ-level masked reconstruction to learn localized anatomical context, (ii) organ-level vision-language alignment to associate organ-specific visual features with organ-name text embeddings, and (iii) organ-level supervision to preserve voxel-level structural detail. SOAR integrates easily with existing LLMs and is evaluated on five public/in-house benchmarks, improving multiple-choice VQA accuracy by 3.53% and disease screening/abnormality detection AUC by 4.25 points on average over all benchmarks/baseline methods. Code will be available.