Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context
Alagappan Valliappan
Abstract
Long-horizon agents are what make long context expensive to serve: a session's KV cache grows turn after turn, and every decode step pays to read it. Frontier LLMs answered that cost on the target side, moving to hybrid and linear attention -- but the built-in Multi-Token-Prediction (MTP/NEXTN) draft heads they shipped with largely kept full attention, and the draft re-reads the cache on every forward. That co-design gap inverts the very optimization it was meant to deliver. One draft layer issues $\gamma$ full-context reads per decode step while a hybrid target issues one per full-attention layer, and only 9-25% of its layers are full attention in the models we test, so the two read counts become comparable. At long context those reads are what set step time. Measured at $\gamma=6$, the entire draft phase adds +92% to +138% to per-decode-step latency on top of the bare verify, and on hard, low-acceptance tasks a deep native draft turns net-negative -- slower than no speculation at all. We give a per-decode-step cost model that predicts this inversion, then close the gap at the serving layer with Windowed-MTP: a StreamingLLM-style window plus attention sink on the draft's attention only, training-free, drop-in, and target-distribution preserving by construction: the draft only proposes, and the full-attention target still decides the output. Across three architecture families at 1M on a single B200 in SGLang it cuts per-decode-step cost by +28% to +44% (input-invariant, widening from 261K to 1M) and turns the 7.7-11.1% of total KV held by the now-dead draft pool into reclaimable serving capacity. On genuinely multi-turn agent traffic (LongMemEval_M, ~1.5M tokens of accumulated dialogue) the margin is wider, not narrower -- and a cost model calibrated on dialogue alone predicts the document-corpus inversions it never saw.
Chat is not available.
Successful Page Load