Training Quality Determines Efficiency Boundaries in Test-Time Reasoning
MD Azizul Hakim
Abstract
Training quality is the primary determinant of reasoning efficiency in large language models, yielding greater returns than proportional increases in inference computation. Evaluating 44 language models (0.5B--685B parameters) across six reasoning benchmarks totalling over 67000 assessments, this study establishes when additional inference tokens are beneficial, redundant or actively harmful. Standard models exhibit near-universal accuracy convergence at approximately 1000 tokens regardless of architecture or parameter count. Causal intervention experiments demonstrate that constraining generation length improves reasoning accuracy by 15.4 percentage points, identifying over-generation rather than capacity exhaustion as the primary efficiency bottleneck. Among 13 reasoning-specialised models, training methodology yields 18.8-fold efficiency improvement at matched parameter scales, exceeding gains from ten-fold parameter increases. Graduate-level evaluation reveals that arithmetic benchmark performance does not predict reasoning capability on complex tasks ($r = -0.03$, $P = 0.91$), exposing fundamental limitations invisible to standard evaluation practice. Mechanistic profiling reveals two orthogonal dimensions of training quality---quality control and decomposition efficiency---that independently predict efficiency profiles, providing a measurable framework linking training decisions to inference-time behaviour.
Chat is not available.
Successful Page Load