Evaluating the Ranking Layer of Rare-Disease Decision Support: Bayesian Uncertainty Under Temporal and Patient Shift
Bonaventure F. P. Dossou ⋅ Mame Diarra Toure
Abstract
Whether a variant ranker transfers across patients, and whether its uncertainty flags unfamiliar genes, are questions that a single solved proband cannot answer. They are also properties of the ranking layer that a generative clinical assistant inherits and cannot verify on its own, because a fluent explanation can make an incorrect ranking appear more credible. We study this gap through the Rare Disease, Real Kid MVA Hackathon 2026 and three post hoc generalization benchmarks. A phenotype-guided whole-genome pipeline processed 5.08 million normalized variants and ranked the accepted compound-heterozygous \bub{} pair first without a disease-specific bonus, achieving 100 Rank Points and F-max 1.0. We compare deterministic baselines, deep ensembles, singular Bayesian neural networks (SBNNs), Bayes-by-Backprop (BBB), and wider Bayesian architectures. On 313,171 temporal \clinvar{} records, a wide rank-50 SBNN achieves the best simulated top-1 ranking (84.35\%) and F-max (0.9235), while the ensemble gives the strongest unseen-gene OOD detection (AUROC 0.7485). On 719 deduplicated Phenopacket cases whose causal variants were absent from ClinVar 2023, HPO conditioning raises dense BBB top-1 from $34.1\%$ to $58.9\%$, but provides no aggregate improvement in the subgroup whose causal genes are absent from the phenotype reference. On 275,739 VarPB spike-in trials across 107 real-genome exonic backgrounds, a predictor-augmented SBNN reaches 86.63\% Top-100, compared with 84.82\% for REVEL and 69.72\% for AlphaMissense. These results motivate reporting molecular scores, model uncertainty and phenotype-reference coverage separately.
Chat is not available.
Successful Page Load