ToLD: Efficient Time Series Forecasting via Tokenized Truncated Latent Diffusion
Abstract
Probabilistic long-term time series forecasting requires estimating a distribution over future trajectories under practical deployment constraints, where inference latency and stability can be as critical as forecast quality. Existing diffusion-based forecasters provide expressive uncertainty modeling, yet their iterative denoising incurs substantial horizon insensitive runtime overhead, which limits real time long-term use. In contrast, one-shot generative models such as VAEs and flows are computationally efficient but often struggle to capture complex multi-modal structure and heavy tailed uncertainty. We propose \textbf{ToLD}, an efficient time‑series forecasting framework built upon \textbf{to}kenized truncated \textbf{l}atent \textbf{d}iffusion to enable fast and reliable long term generation. ToLD adopts a tokenized conditional latent modeling and performs compact token-wise refinement via a few-step truncated latent diffusion process, where the diffusion noise is injected in a context-dependent manner to model non-stationary uncertainty. To reduce train test mismatch in multi-sample forecasting, we further introduce a score distillation scheme that learns a target-free scoring function for test-time candidate ranking. Experiments on eight real-world datasets show that ToLD improves both point accuracy and probabilistic quality over state-of-the-art, achieving up to \textbf{4.3\%} MSE improvement and \textbf{19.3\%} CRPS reduction on long-term tasks with minimal additional inference overhead.