Measuring What Histology Can Predict: Leakage-Stratified Evaluation for Spatial Transcriptomics
Sajib Acharjee Dip ⋅ Liqing Zhang
Abstract
Models that predict spatial gene expression from H\&E histology are judged by one number: the per-gene correlation between predicted and measured expression. We show that number can move by an order of magnitude without touching the model. On $360$ Visium slides across $12$ organs, holding the data, the genes and the predictor fixed, three choices in how the evaluation is built decide the result. The first is what you hold out. Moving from random spot splits to holding out whole studies costs a factor of $1.3$ to $6.7$, and accuracy falls at every step in all $12$ organs. Patient-level splits, the current standard, buy almost nothing, because most slides come from a single donor. The second is whether correlations are pooled across slides: under pooling, a predictor that is told each slide's average expression and nothing else recovers $59$--$99\%$ of the reported score. The third is the missing denominator. We estimate how reliably each gene is measured in the first place, and we also measure that estimator's floor, which turns out to be high enough to matter. Holding out studies also shrinks the training set, so we regroup slides at random into fake studies of the same sizes. Between $58\%$ and $77\%$ of the drop survives, which shows it comes from the grouping and not from having less data. The same pattern appears for four kinds of predictor and three image encoders, including a ResNet-50 that has never seen a tissue slide. We release the protocol as tested software.
Chat is not available.
Successful Page Load