Approaching the Acceptance-Limited Ideal Speedup: Speculative Decoding on Tenstorrent Galaxy
Johanna Rock ⋅ Utku Aydonat ⋅ Grayson Rechsteiner ⋅ Allan Liu ⋅ Saad Jameel ⋅ Daniil Yurshevich ⋅ Maksim Tsishkouski ⋅ Demetris Chrysostomou ⋅ Ambrose Ling ⋅ Adyan Hossain
Abstract
Speculative decoding accelerates large language model generation by drafting several future tokens and verifying them in one target-model pass. Its ideal speedup is bounded by the mean acceptance length, while drafting, multi-token verification, communication, and runtime overhead determine how much of this bound is realized. We study this hardware-software interaction on Tenstorrent Galaxy systems using two structurally different drafters: DeepSeek-R1 multi-token prediction (MTP) and GPT-OSS-120B DFlash. We implement MTP with TT-NN tensor and expert parallelism on four Wormhole Galaxies and with persistent pipeline parallelism on Blackhole Galaxy; DFlash uses pipeline parallelism on Blackhole Galaxy. TT-NN MTP1, pipeline MTP1, pipeline MTP3, and DFlash achieve $1.60\times$, $1.68\times$, $2.00\times$, and $3.57\times$ decode speedups, corresponding to 88.9\%, 93.1\%, 93.0\%, and 87.1\% of their acceptance-limited ideals. In a separate five-point matched-acceptance DFlash study ($\tau\approx2.1$-$5.6$), Blackhole Galaxy at concurrency four achieves $1.37$-$3.28\times$ speedup, versus $1.20$-$1.38\times$ on one H200 at concurrency one, with Galaxy showing higher speedup at every matched point.
Chat is not available.
Successful Page Load