Multi-modal Latent Predictive Architectures for Heterogeneous Clinical Data
Abstract
Multimodal Electronic Health Records (EHRs) combine heterogeneous sources such as billing codes, medication histories, and clinical text. A central challenge is transforming this information into representations where cross-modal synergies are accessible without discarding unique modality signals. Existing generative integration frameworks waste bottleneck capacity reconstructing low-level noise and require complex, heterogeneous loss functions. To address this, we introduce the Multi-modal Latent Predictive Architecture (MLPA), a framework for self-supervised, task-agnostic learning. MLPA combines frozen unimodal encoders, a capacity-constrained Multimodal Bottleneck Transformer, and a uniform masked latent predictive objective. Evaluations on the MIMIC-IV dataset expose a fundamental trade-off between intrinsic representation geometry and predictive accessibility. Pure bottleneck fusion organizes the latent space into highly cohesive clinical clusters but discards unique predictive signals, termed fusion compression loss. Augmenting the compressed bottleneck with direct access to unimodal pathways bypasses this loss, maximizing downstream linear decodability. These findings demonstrate that effective multimodal integration requires simultaneous access to raw modality features and fused structural summaries, framing EHR integration as a problem of targeted information routing and self-supervised compression.