Temporal Structure at Adapter Scale for Longitudinal Brain MRI in Post-Treatment Glioma
Vayangi Ganepola
Abstract
Most brain MRI foundation models (FMs) are pretrained on single volumes, and registration FMs, though trained on scan pairs, optimise alignment rather than disease change, so clinically meaningful change over time is a signal none of them directly rewards. We test this in post-treatment glioma, where separating progression from treatment effect requires comparing consecutive scans and cannot be inherited from competence on static anatomy. Across three public FMs and two cohorts, no FM dominates, from-scratch baselines stay competitive, and which encoder ranks first depends on the readout. We propose T-LoRA, a temporal adapter, as an efficient new direction for adapting brain MRI FMs to disease monitoring. This matters most where resources are scarce. Pretraining an FM is out of reach without cluster-scale infrastructure, and adapting the wrong checkpoint wastes compute that cannot be recovered, so the practical question is not whether an encoder is strong in general but whether it carries the signal a longitudinal task needs, answered before the adaptation budget is spent. We evaluate contrastive (BrainIAC \cite{brainiac}), reconstruction (BrainSegFounder \cite{BrainSegFounder}) and registration (uniGradICON \cite{unigradicon}) encoders across three tasks of rising temporal demand, using UCSF-ALPTDG \cite{ucsf_alptdg} and MU-Glioma-Post \cite{mu-glioma}, neither listed in any encoder's published pretraining data. Heads, losses and schedules are fixed so only the encoder varies; three seeds, one 32\,GB GPU. No FM dominates across tasks and cohorts, the pair-trained registration encoder included: the contrastive encoder leads with global readouts and the reconstruction encoder with dense ones, and in change classification BrainIAC leads on UCSF-ALPTDG (AUROC 0.79) while a from-scratch 3D CNN leads on MU-Glioma-Post (AUROC 0.81). This points to representation-readout alignment, not a temporal prior from pretraining, as what separates the FMs. If the decisive factor is the readout rather than pretraining, temporal structure belongs at adapter scale. T-LoRA keeps the frozen backbone and rank-$r$ bottleneck of LoRA \cite{hu2022lora}, adding two mechanisms in that code space. FiLM modulation \cite{perez2018film} conditioned on the inter-scan interval and each scan's role makes elapsed time and ordering explicit rather than implicit in a difference vector. An $r \times r$ mixing matrix then lets the timepoints exchange information at $r^2$ parameters instead of $d^2$, or 64 per adapted block at rank 8. T-LoRA experiments are in progress. The open question is whether adapter-scale conditioning suffices, or whether change requires objectives that pair scans; this decides whether longitudinal imaging stays gated behind pretraining budgets. Both cohorts arrive skull-stripped, co-registered and complete in all four sequences, and every experiment uses the full training set, so nothing here speaks to missing sequences or the low-label regime where these encoders should help most.
Chat is not available.
Successful Page Load