Windowed-MTP: Removing the Full-Context Draft-KV Tax in Million-Token Agentic Serving
Alagappan Valliappan
Abstract
Agentic workloads make inference stateful: a long-horizon trajectory accumulates context across turns and tool calls until every decode step reads a million-token KV cache. Speculative decoding is the standard latency lever there, and it assumes the draft is negligibly cheap. On hybrid-attention targets that assumption is false, and agentic context lengths are what expose it. This is an unaccounted resource cost with a countable signature: targets moved to sub-quadratic hybrid attention while their built-in Multi-Token-Prediction (MTP/NEXTN) draft heads largely kept full attention, so one draft layer issues $\gamma$ full-context reads per decode step against the target's one per full-attention layer, and only 9-25% of its layers are full attention in the models we test. The draft's context-read work is therefore a fixed fraction of the target's at any context, and that fraction is large. Context length decides only whether the work costs time: at short context a decode step is spread across weight traffic and fixed per-forward costs, whereas at 1M attention dominates it. Measured at $\gamma=6$, the draft phase adds +92% to +138% on top of the bare verify, and on hard, low-acceptance tasks a deep native draft turns net-negative -- slower than no speculation at all. We give a per-decode-step cost model that predicts this inversion, then close the gap at the serving layer with a StreamingLLM-style window plus attention sink on the draft's attention only (Windowed-MTP): training-free, drop-in, and lossless by construction, since the draft only proposes and the full-attention target still decides the output. Across three architecture families at 1M on one B200 in SGLang it cuts per-decode-step cost by +28% to +44% (input-invariant, widening from 261K to 1M) and turns the draft's KV pool (7.7-11.1% of total KV) into reclaimable capacity -- an extra concurrent 1M-token session on the same GPU. On genuinely multi-turn agent traffic (LongMemEval_M, ~1.5M tokens of accumulated dialogue) the margin is wider, not narrower.
Chat is not available.
Successful Page Load