scTrilemma: Balancing Identity, Invariance, and Expression in Single-Cell Representation Learning
Abstract
Single-cell RNA-seq representation learning is inherently label-free, because cell identities, states, and tissue contexts are discovered through analysis rather than provided as training targets. Whether a source of variation is signal or nuisance then depends on the analysis, so the cell embedding faces competing demands, which we call the representation trilemma. It should preserve biological identity and state, remain robust to nuisance context, and retain the gene-level variation needed for expression analysis. We introduce scTrilemma, a latent-bottleneck VAE that treats this tension as an information-routing problem and tests whether expression-derived routing alone can balance the three demands without target annotations or auxiliary representation losses. scTrilemma combines expression-gated gene encoding, cell-representation routing through the decoder, and pseudo-bulk-derived prior conditioning under a single reconstruction objective. In release-based zero-shot evaluation on successive CZ CELLxGENE Census releases, scTrilemma maintains strong biological identity while improving context invariance and expression fidelity, and preserves biological-state, differential-expression, and pathway structure across multiple disease settings.