Reasoning-Aware IRT: Explainable Evaluation of Test-Time Scaling in Reasoning Models
Abstract
Reasoning models increasingly depend on test-time computation, but raw accuracy conflates intrinsic ability, problem difficulty, guessing effects, and reasoning budget. Item Response Theory (IRT) adjusts for problem difficulty, yet collapses test-time compute into a single latent score. We propose reasoning-aware IRT, an extension in which each model's effective ability depends on baseline ability, reasoning-budget responsiveness, and overthinking curvature. Parameter estimation is performed via multi-pass generalized EM with post-hoc initialization refinement. On LiveBench with 42 language models, our method improves rank stability, gives stronger held-out correctness discrimination, and reveals interpretable structure unavailable to classical IRT: a unified benchmark--ability scale with quantified reasoning-token benefits, budget-efficient and overthinking regimes, and rank reversals obscured by raw scores. Our framework offers a compact, explainable evaluation tool for reasoning models under test-time scaling.