PRISM: Pathology Reliability In Scarce-label Medicine
Alicem Koyun
Abstract
Pathology foundation models are typically benchmarked on classification performance at full label availability and on in-distribution test sets, while the properties that govern clinical deployment (calibration, behavior at low label budgets, and robustness under cross-institutional shift) are seldom evaluated jointly. We introduce PRISM, a tile-level benchmark that measures the reliability of pathology foundation models along four axes simultaneously: model choice, label fraction, calibration, and out-of-distribution transfer. PRISM evaluates eight foundation models (CLIP, PLIP, CONCH, UNI, VIRCHOW2, GigaPath, H-Optimus-0, MIDNIGHT) on six tile-level datasets at six label fractions (1% to 100%) with three seeds, 864 linear probe runs in total, extended with a five-hospital Camelyon17 study of 20 directed transfer pairs that holds the label definition fixed. We report AUROC, Expected Calibration Error before and after temperature scaling, Brier score, and a composite Clinical Readiness Index (CRI). Four phenomena emerge. (i) Under full supervision AUROC rank predicts calibration rank consistently across datasets (pooled Spearman $\rho = 0.57$, $I^2 = 0\%$), whereas at 1% labels the relationship is no longer detectable and becomes dataset-specific ($\rho = -0.30$, $I^2 = 67\%$), a significant difference ($\Delta\rho = 0.62$, 95% CI [0.08, 1.13]). (ii) On the harder datasets the linear probe collapses to a single predicted class below a dataset-specific label fraction. (iii) Whether additional source-domain labels help or harm target calibration is governed by how well the probe transfers, not by the presence of shift: below a transfer AUROC of roughly 0.85 more labels worsen target ECE in 86-88% of (model, pair) combinations and above it in 1-8%, a relation that holds within a single dataset and among combinations that never produce degenerate predictions. (iv) Post-hoc temperature scaling effectiveness varies by more than 50x across cells, and a temperature fitted at the source is actively harmful under failing transfer. PRISM is released as a pip-installable package with pre-computed embeddings and reference results for all 288 (model, dataset, label fraction) cells. Code and data: https://anonymous.4open.science/r/prism-benchmark-0108/.
Chat is not available.
Successful Page Load