STEdgeFormer: MCU and NPU Deployable Transformer Architecture
Abstract
Transformer-based sequence models have achieved strong performance in language and time-series modeling, but self-attention remains computationally and memory intensive for deployment on resource-constrained edge accelerators. We propose \textit{STEdgeFormer}, a lightweight, hardware-aware sequence architecture designed for efficient execution on embedded neural accelerators. \textit{STEdgeFormer} combines polynomial mixture layers and DyT normalization with an accelerator-aware causal memory block, replacing pairwise self-attention with compact causal aggregation. On TinyStories, the fully NPU-supported \textit{STEdgeFormer (MW)} achieves a perplexity of 7.39 with 4.12M parameters and 6.51 MB RAM, with an end-to-end latency of 40 ms. In comparison, the standard 4.12M-parameter configuration achieves a lower perplexity of 6.79 but requires CPU fallback for unsupported operators, resulting in 1450 ms latency. Vanilla transformer on the other hand achieves 5.71 perplexity score using 4.40M parameter model but it can't be deployed on the hardware. On downstream tasks, \textit{STEdgeFormer} achieves 94\% accuracy on ECG5000 and 93\% accuracy on Speech Commands with 1.39M parameters. These results demonstrate that hardware-aware architectural choices and operator compatibility can have a substantial impact on practical inference latency, beyond model size and task accuracy alone. Rather than targeting state-of-the-art accuracy on individual benchmarks, \textit{STEdgeFormer} is designed to provide a deployable sequence modeling architecture under the memory, computational, and operator constraints of embedded neural accelerators.