NPCBench: A Clinical Apprenticeship Benchmark for Guideline-Constrained Care-Pathway Reasoning in Nasopharyngeal Carcinoma
Abstract
Current medical LLM benchmarks typically decompose clinical competence into isolated factual questions, single-image interpretation, or brief case scenarios. This fragmented paradigm primarily measures local correctness but provides limited insight into whether models can progressively integrate subspecialty knowledge, multimodal evidence, and evolving patient states into safe, guideline-constrained longitudinal reasoning. To address this gap, we introduce NPCBench, a multimodal clinical apprenticeship benchmark for nasopharyngeal carcinoma (NPC). NPCBench operationalizes specialist training as a staged evaluation of (M)LLMs across guideline recall, atomic rule application, multimodal evidence grounding, longitudinal care planning, state revision, and expert-style consultation. It comprises 4,077 staged evaluation problems and 25 multimodal full-care path episodes, covering pathology diagnosis, MRI staging, and 146 downstream clinical decisions. The benchmark explicitly links stepwise tasks to complete patient-level care trajectories, enabling systematic evaluation of all clinical pathways. Beyond isolated precision, NPCBench evaluates whether the models ground decisions in multimodal evidence, maintain temporally consistent patient states, and provide safe expert-level consultation across complete care trajectories. Across 30 evaluated models, we observe a persistent composition gap: strong performance in local case-level tasks does not translate into episode-level clinical competence required for end-to-end patient management. Overall, NPCBench advances medical AI evaluation from static question answering toward apprenticeship-style workflow evaluation, providing a rigorous testbed for evidence-grounded and temporally coherent reasoning in complex oncology care.