Patient-Held-Out Evaluation of Pathology Foundation Models for Histology-to-Expression Prediction
Abstract
Histology reveals tissue structure, whereas spatial transcriptomics measures gene activity at known tissue locations. Predicting expression from histology could extend molecular analysis to specimens without a spatial assay. Reliable evaluation requires testing on patients excluded from training and distinguishing improvements in prediction from the selection of easier genes. We assemble 144 human tissue sections from five public sources, covering five organs and 79 source-scoped patient identifiers. The collection contains 263,232 paired image and expression spots, with source provenance, patient grouping, and quality eligibility recorded for each section. After quality filtering and a minimum patient-count requirement, 119 sections from 62 patients across four organs form the primary benchmark. We compare five frozen encoders under a common regression protocol with each patient held out in turn. Every tested pathology encoder exceeds the ImageNet baseline in mean predictive correlation in all four cohorts. We then hold predictions fixed and compare gene-selection rules derived from training patients. Selecting 20 of 200 genes by spatial organization, measured using Moran’s I, raises mean correlation in 19 of 20 encoder-cohort settings. However, spatial ranking outperforms ranking by mean expression in brain and breast, while mean-expression ranking performs better in heart and kidney. This ordering is consistent across all five encoders. The results distinguish improvements associated with image representations from those obtained by selecting more predictable targets. The accompanying manifests, patient folds, and evaluation outputs support separate assessment of these effects