Context Windows Are Not Context Capacity: Retention Profiles for Long-Context Foundation Models
Krishna Harish ⋅ Ishant Ghimiray
Abstract
A model that accepts one million tokens does not necessarily use one million tokens reliably. Existing long-context evaluations increasingly expose this gap, but their summaries remain heterogeneous: raw averages mix short-context competence with length robustness, per-length normalized scores require many numbers, and thresholded “effective length” depends on an arbitrary cutoff and is often censored by the largest tested length. We treat long-context ability as a retention profile rather than a scalar window. For model $m$ on benchmark $b$, we normalize the length-conditioned score into $R_{m,b}(L)=S_{m,b}(L)/B_{m,b}$ and introduce log-integrated retention (LIR), context half-lives $H_\tau$, log-length elasticity, and retention dominance. This framework subsumes two existing ideas: 100-LongBench’s LongScore is $R(L)-1$, while NoLiMa’s effective length is $H_{0.85}$. We prove that LIR is stable to uniformly bounded score perturbations whereas threshold lengths can be arbitrarily unstable near flat crossings or become right-censored. We then conduct a reproducible secondary analysis of public RULER (33 models) and NoLiMa (22 models) result tables. On RULER, advertised context length has only moderate descriptive rank association with LIR ($\rho=0.437$); several large-window models retain substantially less capability than shorter-window peers. On five shared model labels and a common 4K–32K support, RULER and NoLiMa retention differ by 0.187–0.496 LIR points, showing that “context capacity” is strongly task-conditioned. We therefore recommend reporting a Context Retention Card—the curve, baseline, LIR, a small family of $H_\tau$ values with censoring, and the largest local drop—instead of a single context-length claim.
Chat is not available.
Successful Page Load