SPICI: Signal Property-Induced Classification and Interpretability for Time-Series Transformers
Abstract
Transformers are increasingly applied to time series classification (TSC). Prior work, TranSPAN(1), showed that in Transformer-based time series classification (TTSC) models, attention heads capture fundamental signal properties that are critical for TSC (frequency, amplitude, phase, trend, seasonality, and sharp transitions), but it remains limited to post-hoc observation, without inducing property specialization during training, nor does it link specific properties to individual class decisions. This paper proposes SPICI (Fig: 1), a framework that actively encourages attention heads in TTSC models to specialize in specific signal properties during training. First, a shared Signal Property Head (SPH) module processes each encoder layer’s output and predicts property scores for the six signal properties. These predictions are supervised by pre-computed ground-truth scores derived from wavelet decomposition (Daubechies-4, level-3). Each layer’s loss is then weighted by the properties it contributes to, based on TranSPAN’s findings, preserving the hierarchical property organization across encoder depth. Second, a head-level entropy loss is applied directly to the distribution of each attention head. For each head, the output is decomposed using the wavelet transform (Daubechies-4, level-3), producing three detail bands (high, mid, low frequency) and one approximation band (lowest frequency). The energy of each band is converted into a probability distribution over the bands. The entropy of this distribution serves as the loss for that head. High entropy indicates an unfocused head; low entropy indicates band specialization. Minimizing this loss drives each head toward frequency-band specialization, which SPH supervision then links to specific signal properties. At inference, the SPH module is removed entirely, so the deployed model is architecturally identical to the original backbone with no inference cost. We evaluate the framework on six TTSC architectures (vanilla Transformer, TrajFormer, SVP-T, MTM, TARNet, GTN) across three benchmark datasets (ECG5000, FordA, UCI HAR), following TranSPAN’s setup. The framework improves accuracy across all model-dataset combinations (average +2.1 points), with GTN reaching 95.0–96.2% (up from 93.1–94.0%) and the vanilla Transformer gaining the least (+0.8 to +1.7). The number of disentangled attention heads, those specialized for exactly one signal property, increases substantially across models, with GTN, MTM, and TARNet each approximately tripling their counts (from 5–8 to 20–22 across all three datasets), SVP-T reaching the highest count of any model (31 on FordA), and the vanilla Transformer moving from zero to 6–9. Using a class-flip protocol, we mask attention heads attributed to a single signal property and measure the fraction of correctly classified instances whose prediction changes (flip rate). Across all three datasets, flip rates of 3.9–8.8% against a 1.0% matched-random control reveal that transient classes (Anomaly on FordA, R-on-T PVC, PVC on ECG5000, Walking on UCI HAR) depend on sharp transitions and frequency, while baseline classes (Normal on FordA and ECG5000, static postures on UCI HAR) depend on trend, amplitude, and phase. Seasonality is active only on UCI HAR, matching the periodic nature of gait. We prune heads not responsible for any property while retaining disentangled heads. GTN tolerates 26.6–34.8% pruning across all three datasets around 2.3 percentage points accuracy loss, and SVP-T tolerates 23.3–34.5% with nearly 1.7 points of accuracy loss, without fine-tuning. Critically, removing property-specific heads selectively damages the dependent classes, while size-matched random removals damage both classes equally across all datasets, causally validating the disentanglement claim. These findings offer a generalizable pathway for achieving structured interpretability and compression in TTSC models without architectural modification. The resulting property-level head specialization further provides a foundation for developing explainable time series architectures. References [1] N. N. Munasinghe, S. Wickramanayake, and D. Meedeniya, “Transpan: Analyzing signal property learning in transformer-based time series classification,” IEEE Access, vol. 14, pp. 77 776–77 790, 2026. Figure 1: Overview of the SPICI framework.