Individualized and Subtype-Aware Representation Learning in Biomedical Data
Abstract
Machine learning is increasingly used to uncover complex disease patterns in biomedical data and to support personalized diagnosis, prognosis, and treatment planning. However, most existing models fail to capture the substantial heterogeneity underlying complex diseases. Patients diagnosed with the same condition often exhibit diverse molecular profiles, progression trajectories, treatment responses, and latent disease mechanisms. This challenge is amplified in biomedical datasets characterized by limited sample sizes, weak supervision, batch effects, and high-dimensional noise. As a result, standard deep learning approaches tend to learn representations dominated by shared background structure, obscuring subtle but clinically meaningful disease-specific variations. Our work is motivated by the observation that many biomedical conditions are not monolithic categories, but collections of heterogeneous latent subtypes that manifest differently across individuals. Existing classification-oriented approaches often impose rigid diagnostic boundaries that fail to reflect this underlying pathological heterogeneity. In contrast, CASL-VAE reframes biomedical representation learning as a personalized latent modeling problem, where the objective is not only to distinguish diseased from healthy populations, but also to characterize how individual patients deviate from shared normative structure and whether these deviations are systematically shared across subpopulations. This formulation enables the discovery of clinically meaningful subtype organization and patient-specific disease patterns that may remain hidden in conventional predictive frameworks. To this end, we introduce CASL-VAE, a structured contrastive latent variable framework for learning interpretable representations from unpaired reference (healthy) and target (disease) biomedical data. CASL-VAE decomposes variation into (i) a common latent space that captures shared structure across populations and (ii) a hierarchical salient latent space that models disease-specific subtype structure and within-subtype variability. The framework combines variational inference with structured regularization to encourage separation of shared and target-specific variation while enabling unsupervised subtype discovery and paired-sample generation which inturn supports cross-domain comparisons. We evaluate CASL-VAE across semi-synthetic neuroimaging data, real Alzheimer's disease cohorts, transcriptomic data, and synthetic image datasets. In semi-synthetic experiments, CASL-VAE outperforms existing contrastive analysis and generative modeling baselines in subtype discovery, particularly in challenging settings with subtle disease effects. On real Alzheimer’s disease data, the model identifies two biologically plausible neuroanatomical atrophy patterns consistent with known signatures of Alzheimer’s pathology and produces participant-level deviation indices associated with cognitive and biomarker variables. Notably, CASL-VAE demonstrates stable performance across modalities using the same model configuration without dataset-specific tuning. This work contributes to the intersection of generative modeling, contrastive learning, and precision medicine by introducing a structured latent framework for modeling heterogeneous variation in unpaired biomedical data. CASL-VAE provides a flexible and interpretable approach for subtype discovery, individualized disease characterization, and exploratory analysis in complex biomedical settings.