scTrilemma: Balancing Identity, Invariance, and Reconstruction in Single-Cell Representation Learning
Abstract
Single-cell RNA-seq representation learning is inherently label-free: cell identities, states, and tissue contexts are typically discovered through analysis rather than provided as target labels during training. As a result, the evaluated cell embedding must balance competing demands, which we refer to as the representation trilemma: it should preserve biological identity, avoid encoding collection effects as cell identity, and retain expression information needed for reconstruction. We introduce scTrilemma, a latent-bottleneck VAE that tests whether simple expression-derived routing can manage this trilemma without cell-level annotations, metadata-derived supervision, auxiliary representation losses, or specialized disentanglement modules. scTrilemma combines expression-gated gene encoding, cell-representation routing through the decoder, and pseudo-bulk prior conditioning under a single reconstruction objective. In release-based zero-shot evaluation on successive CZ CELLxGENE Census releases, scTrilemma improves batch-effect removal while preserving biological identity, retains marker-program and biologically meaningful tissue-context structure in the evaluated embedding, and better preserves differential-expression and pathway structure through reconstruction; ablations support burden allocation as the source of these gains.