Sparse Expansion Utility: Identifying and Routing to Pivotal Steps in LLM Reasoning Chains
Abstract
LLM chain-of-thought reasoning improves with additional compute, but standard 2 inference scaling [Wei et al., 2022, Wang et al., 2022, Snell et al., 2024, Brown 3 et al., 2024] treats all steps equally. We ask: which step in a reasoning chain benefits 4 most from extra computation? We make three layered contributions, each scoped 5 explicitly. (I) Measurement. We define expansion utility Uk as the marginal 6 gain in correctness probability from resampling from step k onward, measured via 7 a step-oracle sweep. Across nine model families (8B–123B), seven benchmarks 8 (math, science, code), and 6,500 chains, every one of the 41 healthy (model, bench) 9 cells shows positive oracle gap at 10% step budget (95% Wilson CI [0.914, 1.000]; 10 mean +22.0 pp, range +1 to +48 pp). (II) Structure. Per-step utility is sparse 11 (97.6% of cells: Gini ≥ 0.80); Gini and mean oracle gap are tightly anti-correlated (Pearson r = −0.913, p = 9.5 × 10−17 12 ). A resample-stability simulation and 13 a segmentation-rule sensitivity check (Appendix J) bound the headline against 14 noise and rule-choice confounders. (III) Routing proof-of-concept. A cross15 model router trained with a within-example listwise loss [Cao et al., 2007] and a 16 learned depth prior, combined at inference with a bench-conditional depth mask 17 and a confidence gate, recovers 5-seed-averaged gated GR@10% = +0.149 (95% 18 bootstrap CI [+0.085, +0.217]) of the cross-model oracle gap at matched step19 budget, with no architectural change to the underlying model. We frame this as 20 proof the framework is actionable cross-model, not as deployment-ready scaling; 21 matched-token comparison to self-consistency / best-of-N requires a separate 22 generation campaign and is left to follow-up (§7). Methodological observation. 23 Within-example pointwise top-k BCE under cross-model pooling produces a high24 AUC discriminator (AUC@10 = 0.70) that selects the wrong steps (gated GR 25 = −0.15); the listwise loss is shift-invariant to per-(model, bench) utility scale and 26 removes this inversion.