PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us
Gal Sapir ⋅ Alon Diament ⋅ Adva Wolf ⋅ Doron Yaya-Stupp ⋅ Dikla Gelbard-Solodkin ⋅ Dana Azouri ⋅ Anat Etzion-Fuchs ⋅ Guy Lutsker ⋅ Eran Segal ⋅ Hagai Rossman
Abstract
Deeply phenotyped cohorts measure clinical, imaging, molecular, and wearable modalities in the same participants, combining dense observations across timescales from seconds to days with longitudinal follow-up over years. This breadth can reveal which measurements inform which health-related questions, but results from heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more than 13,000 participants have completed the initial visit. Each PhenoBench question fixes the target, eligible population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. The current benchmark defines 90 clinically grounded tasks across 15 domains and 26 information sources. Measurements showed question- and representation-dependent predictive value, including positive, null, and negative associations. As one controlled demonstration, we used PhenoBench to evaluate emerging tabular foundation models across 160 matched regression comparisons spanning 52 tasks. These models ranked above standard task-specific models in aggregate, but improved on ridge by a median of only 0.004 $R^2$ (95% CI, 0.002–0.007). We then used the same cohort data and evaluation contracts to test 14 language models on 35 PhenoBench tasks spanning phenotype recovery, classification, follow-up forecasting, and participant ordering. Without cohort-specific fitting, language models made informative predictions on some tasks, but showed task-specific capability gaps, shared failures of scale, and rarely surpassed models fitted on the same fields. PhenoBench turns a multimodal longitudinal cohort into a versioned, auditable evaluation system where new questions, measurements, and models can be added without redefining existing comparisons.
Chat is not available.
Successful Page Load