The Committed Operating Point: Decision-Utility Collapse in Continual Multimodal Breast Imaging
Rishabh Jha ⋅ Amrita Singh
Abstract
A breast imaging model is not deployed as a ranking of cases but as a rule: choose a threshold on validation data, write it into a protocol, and use it. Continual learning is evaluated almost entirely with accuracy-based metrics, and none of them is computed at that threshold. The underlying problem is a mismatch of lifetimes: the threshold is committed once and never revisited, while the encoder beneath it keeps moving and the data that would justify a new one has been deleted. We train a frozen BiomedCLIP tower with rank-8 LoRA adapters across histology, ultrasound and mammography, deleting each modality's images when its turn ends. Measuring both families on the same runs, we find forgetting stays between $0.09$ and $0.16$ accuracy while sensitivity at the committed threshold falls by $0.21$ to $0.31$. Repairing the threshold afterwards from what was retained fails a criterion registered before the experiment: cosine fidelity reaches $0.65$ where $0.90$ was required, and does not improve as the budget grows. CALIBRE (Calibration-Aware Low-rank Incremental adaptation with Bounded REtention) therefore keeps a 21\,KB random projection of each deleted modality and scores it with the \emph{current} head at every later step, which makes the committed sensitivity and specificity a differentiable term in the loss rather than something to repair afterwards. It improves balanced accuracy at that threshold by $0.088$ over sequential LoRA and $0.114$ over an EWC-style penalty, and it is the only sequential method that keeps the threshold usable on an external cohort, holding specificity at $0.194$ on mammography where those two baselines fall to $0.065$ and $0.007$. Which of its components delivers that gain is left open at four seeds, and we report the ablations in full.
Chat is not available.
Successful Page Load