Capability Retention under Post-Training Linearization of Small Language Models
Anamaria Roberta Hartl ⋅ Michael List ⋅ Balint Laszlo Szarvas ⋅ Lukas Hauzenberger ⋅ Niklas Schmidinger ⋅ Sebastian Böck ⋅ Günter Klambauer ⋅ Sepp Hochreiter
Abstract
Post-training architectural distillation can replace full self-attention with subquadratic sequence mixers, reducing decoding costs while reusing pretrained weights. Yet its effectiveness for compact Small Language Models (SLMs) remains unclear. We distill parameter-matched mLSTM-SWA students from Qwen3 and SmolLM3 teachers spanning 0.6B-4B parameters and evaluate coding, mathematics, STEM, and tool use. Retention is domain-dependent: within Qwen3, coding and mathematics retention improves with scale, whereas tool-use retention is non-monotonic and stateful multi-turn performance remains near zero at every scale. We then merge independently specialized students to test whether skills can be combined without additional inference-time parameters. Merging reveals a clearer capacity bottleneck: it is near-neutral or beneficial at 3-4B, mixed at 1.7B, and broadly harmful at 0.6B, with tool use degrading at every scale. At long generations, hybrids achieve 3.2-7.1$\times$ higher throughput. Overall, linearization improves long-context efficiency, but expert composition and stateful tool use remain key bottlenecks.
Chat is not available.
Successful Page Load