When Linear-Time Models Look Sublinear: Auditing Scaling Exponents in Long-Context Efficiency Benchmarks
Abstract
Linear asymptotic complexity does not by itself explain latency in the finite long- context ranges used for benchmarking. We introduce the window-sensitivity profile, α(Nmin), which recalibrates the empirical power-law exponent after progressively excluding the shortest measurements. For a four-layer volumetric Mamba en- coder on an RTX A6000, a single resolution sweep of the same encoder design yields α = 0.496 over 1.2 k–76.8 k visual tokens, 0.913 over 9.6 k–76.8 k, and a final two-point secant of 0.956. The resulting 0.460 fitting-window shift is substantially larger than the repeat-level exponent dispersion observed across the three benchmark repeats. Measured latency is nearly flat over the shortest condi- tions before becoming increasingly token-proportional at larger conditions, while profiler-recognized FLOPs exhibit an empirical exponent of α = 1.000 over the measured range. Within this resolution sweep, reducing the visual-token count eightfold, from 9.6 k to 1.2 k, improves measured step latency by only 1.216×. A single fitted exponent can therefore obscure substantial finite-range variation in the measured cost behavior from which it is estimated. We recommend reporting the complete cost curve, its window-sensitivity profile, repeat-level dispersion, and fixed and marginal costs.