Generate Features, Not Frames: Low-Compute Synthetic Priors for Echocardiography Transfer
Harry Mallinson ⋅ River Jiang ⋅ Nicolas Aragon ⋅ ⋅ Puya Vahabi
Abstract
Synthetic echocardiography usually models pixels, even though many downstream tasks use only a learned representation, often called a latent. We generate that latent directly. A conditional rectified flow trained on 1,441 MIMIC-IV-ECHO clips samples 512-dimensional EchoPrime latents conditioned on left-ventricular ejection fraction (LVEF). We test a preselected transfer strategy on two external adult test sets, giving it labels from exactly $N$ target patients and a fixed bank of 10,000 synthetic latents with broad LVEF coverage. Averaged over $N\in\{1,2,5,10,15,20,25\}$, the strategy achieves MAE 6.13 on EchoNet-Dynamic and 9.38 on CAMUS, compared with 9.22 and 11.00 for RidgeCV fitted to the same $N$ labelled patients. By comparison, the same strategy initialized from retained real MIMIC latents performs similarly on EchoNet and worse on CAMUS. Matched-count controls show that Gaussian jitter and latent interpolation do not reproduce the gain. Benefits differ across patient groups: reduced-LVEF error falls on EchoNet but rises on CAMUS. On a pediatric dataset, the adult-selected strategy is worse overall than target-only training but helps the small reduced-LVEF group; adapting sooner recovers much of the overall loss. While this is not a formal privacy guarantee, no generated latent in our audit exactly matched a training latent, and the generated bank provided no measurable additional evidence that any patient was in the generator's training set. Though not suitable for tasks requiring viewable synthetic echocardiograms, task-conditioned latent synthesis can provide low-compute training data when local labels are scarce, offering a practical alternative to video synthesis for cross-site prediction.
Chat is not available.
Successful Page Load