pyCD: A Unified Benchmark for Reliable Evaluation of Cognitive Diagnosis Models
Youheng Bai ⋅ Xueyi Li ⋅ Tengteng Cheng ⋅ Mingliang Hou ⋅ Teng Guo ⋅ Jiaqi Zheng ⋅ Zitao Liu
Abstract
Cognitive diagnosis (CD) infers students' mastery of knowledge components (KCs) from their responses, supporting personalized feedback and exercise recommendation. Cognitive diagnosis models (CDMs) are often ranked by response-prediction AUC under inconsistent experimental protocols, while their KC-level mastery profiles are not consistently audited under a shared protocol. We introduce \texttt{pyCD}, a unified benchmark of eleven CDMs on six public datasets that fixes preprocessing, splits, hyperparameter search, and stopping rules; the benchmark reports KC-level diagnostic metrics for learned mastery alongside standard prediction metrics. Two findings show that gains in AUC do not translate into gains in diagnostic quality. First, the AUC leaderboard is sharply compressed and sensitive to protocol choices: modern CDMs often cluster within $0.005$--$0.020$ AUC, with the per-dataset best non-IRT CDM improving over IRT by only $0.017$ AUC on average; removing validation-AUC early stopping can shift AUC by up to $0.13$ on individual cells. Second, the AUC gaps that remain do not reliably transfer into diagnosis: the top-1 AUC model is typically outside the top-3 by KC-level diagnostic agreement (DOA). The same mismatch extends to exercise difficulty: against Junyi's human pairwise difficulty ratings, learned CDM difficulty parameters rank worse than both an empirical accuracy-based baseline and zero-shot LLMs given only exercise titles. CDM evaluations should therefore report diagnostic outputs directly rather than infer them from predictive gains; \texttt{pyCD} provides a unified protocol for doing so. Code, splits, and configurations are available at \url{https://anonymous.4open.science/r/pyCD-NeurIPS2026-60B8}.
Chat is not available.
Successful Page Load