On-Policy Hindsight Distillation for Early Risk Prediction
Abstract
Early prediction of chronic disease risk from electronic health records (EHRs) is challenging because early clinical signals are often sparse, noisy, and indirectly related to future outcomes. Existing methods typically learn from prediction-time records and final outcome labels alone, which provides limited supervision for identifying weak early signals from heterogeneous and confounded clinical records. LLM-generated rationales offer an intermediate form of prediction-time reasoning, but rationales derived from early EHR may still miss weak signals whose relevance becomes clearer only in later records. We therefore use follow-up EHR observed after the prediction time but before outcome assessment as training-time hindsight and propose On-Policy Hindsight Distillation (OPHD), a self-distillation framework that converts follow-up EHR into training signals for LLM-generated prediction-time rationales. Specifically, OPHD first uses an LLM to generate a rationale from prediction-time EHR alone, then re-evaluates the same rationale under privileged follow-up views to derive token-level signals that are distilled back to the prediction-time policy. To handle noisy follow-up records, OPHD emphasizes hindsight signals that are stable across perturbed privileged views. A task-grounded scorer further prioritizes rationales that improve downstream risk discrimination. Experiments on multiple real-world neurodegenerative disease cohorts show that OPHD achieves the best long-horizon prediction performance compared with a variety of baselines.