Anatomically Aligned Multimodal Diffusion for Individualized Alzheimer’s Disease Trajectory Forecasting
Abstract
Alzheimer’s disease progression models should act as patient-specific trajectory simulators, forecasting which brain regions may change and generating the corresponding future anatomy over a specified interval. Most approaches instead predict diagnostic or cognitive outcomes, while generative models often require paired longitudinal scans and offer limited anatomical control. Multimodal methods also assume synchronized MRI, cognitive, and biofluid measurements, although clinical histories are typically asynchronous and incomplete. We propose an anatomically aligned multimodal diffusion framework in which brain anatomy provides the foundation for multimodal integration and forecasting. An MRI encoder and diffusion U-Net form a diffusion autoencoder that separates the baseline representation into identity-preserving and progression-related components. Mask-guided attention further organizes the progression component into region-indexed anatomical tokens representing the hippocampus, entorhinal cortex, amygdala, ventricles, cortical regions, and other structures. This enables progression-sensitive regions to evolve while preserving patient-specific anatomy. Cognitive assessments and biofluid biomarkers are encoded independently with their acquisition times relative to the baseline MRI. A sparse cross-modal alignment associates each non-imaging token with one or more anatomical tokens before fusion. Clinically supported cognitive-anatomical relationships, such as episodic memory with hippocampal and entorhinal representations, are incorporated as soft priors. Biofluid biomarkers instead receive learned, distributed associations because they are not localized to individual regions. Alignment losses are applied only to observed modalities. An availability mask and structured modality dropout allow missing latent tokens to be estimated from the aligned anatomy and other available evidence without imputing raw measurements. At inference, the model receives only a baseline MRI, the prediction interval, and available time-stamped cognitive or biofluid measurements; future diagnosis is not supplied. The aligned tokens are fused into a patient-specific disease representation, and a time-conditioned trajectory network predicts the future anatomical tokens. These tokens are recombined with the identity-preserving representation and passed to the same diffusion U-Net decoder, generating an individualized next-point MRI without a separate forecasting diffusion model. Future anatomy is the primary output, while auxiliary heads estimate cognitive and biofluid trajectories. Regional-change maps, uncertainty estimates, and trajectory-derived conversion risk support interpretable and personalized disease monitoring. Overall, this work advances Alzheimer’s disease modeling from diagnosis-centered prediction to anatomically grounded, missingness-aware multimodal forecasting of patient-specific future brain states.