Seeing Is Not Knowing: Current Limits of Virtual Staining for Spatial Transcriptomics from Histology
Abstract
Multimodal foundation models are increasingly used to infer spatial molecular state from routine histology. In computational pathology, this has motivated H&E-to-spatial-transcriptomics models that attempt to recover gene-expression maps from hematoxylin and eosin stained tissue sections. We argue that this task is often over-framed as deterministic image-to-molecule translation, despite the fact that morphology only partially observes molecular state. This creates a mismatch between task formulation, model design, and evaluation. First, the H&E-to-transcriptomics mapping is biologically non-identifiable in general: visually similar tissue regions may arise from distinct transcriptional programs, cellular mixtures, or microenvironmental states. Second, spatial transcriptomics platforms define different observation models, changing the granularity, sparsity, noise structure, and gene coverage of the target signal. Third, common metrics such as gene-wise correlation can reward spatial smoothing, tissue-compartment priors, and plausible hallucination rather than specimen-specific molecular inference. We propose failure-aware evaluation protocols based on smoothing controls, cross-cohort and cross-platform stress tests, and uncertainty calibration under morphologically ambiguous inputs. Our goal is to reframe H&E-to-spatial transcriptomics as uncertainty-aware multimodal inference rather than direct molecular recovery from pixels.