Tokenization and Architecture Jointly Allocate Component Roles in Time Series Transformers
Mingyu Kim ⋅ Doguk Kim
Abstract
Time series transformers have achieved strong forecasting performance, yet fundamental questions about their internal mechanisms remain unanswered: among self-attention and feed-forward networks (FFN), which component constructs temporal features? Why does a simple linear model outperform transformers? Why does patching help? We propose a temporal probing framework consisting of 15 token-level diagnostic tasks to trace information flow across transformer components. Applied to 5 architectures across 8 benchmarks, we find that temporal feature construction is not inherently tied to any single component—it emerges from the interaction between tokenization strategy and architectural design. In PatchTST, FFN dominates feature construction across all 8 datasets, with a $3.2\times$ mean ratio over attention. Even when attention is completely removed, FFN's feature construction capability is preserved or enhanced, indicating that attention is not essential for temporal feature construction. Conversely, iTransformer's variate-level attention dominates over FFN, and CATS's cross-attention is responsible for nearly all feature transfer. Point-wise tokenization ($P=1$) structurally deactivates FFN, explaining why early transformers lagged behind linear models. These findings unify contradictory results across DLinear, PatchTST, attention-free models (TSMixer, PatchMLP), and CATS within a single framework.
Chat is not available.
Successful Page Load