SCHTs: A Semi-Structured Dynamic Sparse Training Framework for Hardware-Efficient Deep Learning
Abstract
Although Dynamic Sparse Training (DST) has emerged as a promising solution for deploying deep neural networks on resource-constrained edge devices, unstructured DST methods suffer from irregular weight distributions that incur extra index overhead and severely degrade computational efficiency. In contrast, semi-structured sparsity offers computational efficiency but typically suffers from performance degradation at high sparsity levels with strict semi-structured constraints. To address these challenges, we propose Semi-structured Cannistraci-Hebb Training soft rule (SCHTs), a novel hardware-efficient dynamic sparse training framework, which includes four key stages: semi-structured sparse topological initialization, weight initialization, global network pruning, and local network regrowth. Unlike unstructured DST methods, SCHTs supports fine-grained and coarse-grained N:M semi-structured sparsity patterns with constant fan-in or fan-out, ensuring computational efficiency. Extensive experiments demonstrate that the proposed SCHTs framework enables highly sparse networks (>90\%) to achieve performance comparable to, or even surpassing, their dense counterparts. (1) For spiking neural networks, models trained with SCHTs consistently outperform their dense counterparts at 90\% sparsity, yielding accuracy improvements of +0.05\% (to 94.79\%), +2.95\% (to 75.01\%), and +0.09\% (to 99.16\%) on CIFAR-10, CIFAR-100, and N-MNIST, respectively. (2) For artificial neural networks, sparse networks trained with SCHTs outperform their dense counterparts with accuracy improvements of +0.90\% (to 77.54\%) and +0.63\% (to 63.87\%) using GoogLeNet on CIFAR-100 and TinyImageNet, as well as a +0.67\% (to 78.97\%) improvement using ResNet-152 on CIFAR-100. (3) For large language models, SCHTs enables a 90\% sparse LLaMA-130M to yield a highly competitive validation perplexity of 25.49 on OpenWebText, outperforming unstructured DST methods like RigL. Furthermore, cycle-accurate simulations on Trapezoid (a specialized sparse matrix accelerator) reveal that SCHTs achieves a 75\% reduction in index overhead and a 32.7\% improvement in computational efficiency at 1:16 sparsity, demonstrating its effectiveness in translating algorithmic sparsity into tangible on-chip acceleration.