From Longitudinal Notes to Clinical Trajectories: A Guideline-Grounded Representation for Clinicians and Language Models
Panav Shah ⋅ Pradyumna Chari
Abstract
Guidelines describe what care should look like; longitudinal notes record what care actually looked like. Neither is easy to read at the level a clinician or a model needs, because the same decision is spread across many notes, repeated in different language, and mixed with detail that does not define the pathway. We study an explicit intermediate representation that sits between them: a patient trajectory over a guideline-defined graph of named clinical states and transitions, compact enough to be read by a person and used by a model, and preserving order and branch structure. We formalize the NCCN Meningioma guideline as a graph with 41 states and 58 directed transitions, then use a five-pass LLM procedure to map longitudinal notes onto trajectories. Because every quantity lives on a named guideline edge, one object supports an inspectable per-patient view, cohort pathway maps that make practice variation structurally visible, and a patient-conditioned estimate (a similarity-weighted prior in the power-prior sense) with an explicit measure of the historical evidence behind it. On 250 synthetic patients, scored against the trajectory retained by the data-generating process using a similarity profile that excludes the treatment being predicted, the estimator reaches mean $0.883$ KL divergence over the edge set (the metric that carries transition structure) and $0.780$ Wasserstein distance on node marginals, both better than an unweighted cohort mean. We then ask whether a representation built to be read is also a good teacher. Holding patients, questions, target-writing model, student model, and training settings fixed, a Qwen2.5-0.5B-Instruct model post-trained on trajectory-derived supervision reaches $50.8\%$ accuracy on held-out clinical questions across three seeds, against $43.6\%$ for record-derived supervision, $36.2\%$ for a compact free-text summary written to roughly twice the trajectory's token budget, and $31.7\%$ without post-training. The trajectory beats the record by $7.2$ points from a source $335\times$ smaller, and beats that summary by $14.6$; compression alone does not explain the effect, since the summary is $7.4$ points \emph{worse} than the record. These results support explicit clinical trajectories as a shared representation for pathway analysis and model post-training in a controlled synthetic setting; they do not establish clinical validity on real electronic health records.
Chat is not available.
Successful Page Load