TraceDx: Criticality-Weighted Atomic Facts as a Training Signal for Sequential Clinical Diagnosis
Abstract
In clinical decision-making, success depends on taking the right action to obtain the evidence that matters. One targeted clinical workup can identify a diagnosis that a battery of nonspecific tests missed. Large language models have shown considerable promise in clinical diagnosis, but most current approaches assume that all relevant patient information is available upfront, which does not reflect how diagnostic evidence is gathered in practice. Even when models gather evidence iteratively, challenges arise in determining which action to take, as they receive no feedback on which findings are diagnostically decisive. To solve this problem, we introduce TraceDx, a reinforcement learning framework for sequential clinical diagnosis. TraceDx is built on the idea that unstructured clinical records can be decomposed into atomic clinical findings, each scored by its diagnostic value. This signal rewards agents for discovering decisive evidence in addition to reaching the correct diagnosis. Consistent with this design, a case-level analysis shows that recovery of critical evidence is the only statistically significant predictor of diagnostic success. Direct case comparisons further show that TraceDx uncovers more diagnostically important evidence, and uses it to resolve the key diagnostic uncertainty. TraceDx-trained open-weight models outperform multiple frontier baselines on both MIMIC-CDM and a rare gastrointestinal disease dataset.