When Speculative Decoding Stops Paying Off at Long Context: A Diagnostic Study with RASD
Aman Kesarwani
Abstract
Speculative decoding's 2-3$\times$ short-context speedups do not reliably transfer to million-token inference. This paper asks when and why the transfer stops. On Llama-2-7B extended to 1M tokens with YaRN factor 256, per-round acceptance becomes bimodal: across the measured long contexts, 48% of verify rounds fully reject the draft (pooled Wilson 95% CI $[0.39, 0.57]$) and the median per-round $\alpha$ is exactly 0 at 128k and 512k. The paired per-seed speedup against a matched target-only baseline regresses from $1.76\times$ $[1.59, 1.93]$ at 256k to $0.94\times$ $[0.57, 1.30]$ at 512k and $0.97\times$ $[0.80, 1.14]$ at 1M, statistical parity. A PG-19 dose-response ($\alpha$: $0.71$ at 4k $\to$ $0.26$ at 8k $\to$ $0.11$ at 1M) supports context length as the primary driver. A profiler measurement shows communication is $\leq 1.2\%$ of wall on the speculator rank, ruling out communication-overlap masking. The measured speedup is instead set by a race between two mechanisms: the acceptance collapse and the growing cost of the speculative round itself. Arriving at these findings required a system that runs a target-model forward at 1M context: RASD combines ring-attention sequence parallelism, NF4 KV-cache quantization, YaRN scaling, and speculative verification, and reaches 1M tokens at 39.3 GiB peak per rank on $8\times$ A100, where the vanilla HuggingFace FlashAttention-2 stack runs out of memory at 128k. Code, traces, and the ablation grid will be released upon acceptance.
Chat is not available.
Successful Page Load