Speculative Decoding Is Per-Token, Not Path
Abstract
We prove that the speedup of exact speculative decoding is governed by a per-token rate, not by the path-level prefix-overlap capacity that folklore suggests, and that the gap between the two can be unbounded. The two candidates are the per-token Leviathan rate σL(P, Q) and the path-level coupling capacity ΩL(P, Q) = Σ_{t=1}^L ov(P_{≤t}, Q_{≤t}). We show σL ≤ ΩL always, the inequality is strict already on i.i.d. products, and there exists an explicit fixed-precision binary trajectory-exact rejection-style protocol on L = 2 with E[τ] = 19/16 > 9/8 = σL, ruling out σL as a universal upper bound across all trajectory-exact rounds. Inside the deployable admissible Leviathan-style clipped-ratio class, σL is the tight rate, and the speculation decoding complexity satisfies SDC{B,L}^Q(M; u, T) = Θ(T / (σ_{L,B}(PM^{u,T}, Q) + 1)), achieved by iterated Leviathan and matched up to constants by Wald. On the path side, ΩL exhibits a sharp two-sided dichotomy in the prefix-KL profile ηt: Σt √(ηt / 2) = o(L) forces ΩL = L − o(L), while geometrically bounded overlap ot ≤ ρ^t forces ΩL ≤ 1/(1−ρ) uniformly in L, with no unconditional KL converse. Two explicit fixed-precision Transformer families of size O(log L) realize the two extremes under the same three-element proposal family Q = {Qflat, Qconst0, Qconst_1}: a constant-readout family F⁺ has SDC ≤ T / ((1−ε)L + 1), while an alternating-parity family F⁻ has SDC ≥ T/5, an unconditional Θ(L) architecture separation. Speculative decoding is not path coupling; it is per-token clipped-ratio acceptance, and the architecture of the target controls the rate.