NNCoxKL: Risk-Set Distillation from Probability-Free Prognostic Teachers for Deep Cox Models
Abstract
Survival models are often trained at target clinical sites where outcomes are limited and censored, while patient-level data from external cohorts may be unavailable because of institutional data-sharing constraints. In many applications, the transferable knowledge is not a dataset, a calibrated survival curve, or estimated regression coefficients, but a probability-free prognostic artifact such as a clinician-defined scorecard, guideline-based staging rule, registry-derived risk index, or coarse clinical risk group. These weak teachers may encode human expertise or evidence from prior cohorts, yet they often have only ordinal meaning and need not share a numerical scale with the target model. We propose NNCoxKL, a transfer-learning framework that trains a deep Cox model on target time-to-event data while distilling such human- or model-derived prognostic signals through Cox risk sets. For each event risk set, NNCoxKL converts both the external signal and the neural Cox risk score into Plackett--Luce distributions and regularizes the Cox partial likelihood through a temperature-scaled KL alignment term. This transforms clinical knowledge into a censoring-aware ranking prior over who is most likely to fail next, rather than treating the external score as a covariate or requiring it to represent calibrated survival probabilities. The external signal enters only through this risk-set regularizer, allowing the model to learn nonlinear target-cohort effects while borrowing relative-risk information from external clinical knowledge. Experiments on independent clinical transfer tasks and controlled survival benchmarks show that NNCoxKL improves discrimination and calibration-sensitive prediction over internal-only deep Cox models and common score-transfer baselines.