Attention Architecture and Positional Retrieval in Small Language Models
Abstract
Small, edge-deployable language models (under 1B parameters) are frequently the only viable option in bandwidth- and compute-constrained settings, where reliable cloud access cannot be assumed. Whether these models can actually use information placed in the middle of their context window, rather than only at its beginning or end, is poorly characterized at this scale. We run a controlled needle-in-a-haystack (NIAH) evaluation across four sub-billion-to-1B-parameter instruction-tuned models spanning three attention mechanisms: sliding-window attention (SWA; Gemma 3 270M and Gemma 3 1B), multi-head attention (MHA; Qwen1.5-0.5B), and grouped-query attention (GQA; Qwen2.5-0.5B). Holding evaluation protocol and hardware fixed, we find a stark divergence: Gemma 3 270M’s middle-position accuracy collapses to 0% by 4K tokens and remains at 0% through 16K, while both the MHA and GQA models retain 100% middle- position accuracy at every tested length through 8K. Gemma 3 1B, the same architecture at larger scale, shows a different and non-monotonic failure pattern at 16K, complicating a simple “attention type alone determines severity” narrative. We report our full protocol, directly measured results, and an explicit, clearly- labeled discussion of expected trends beyond our tested range. For practitioners deploying small models in resource-constrained regions with no cloud fallback, these results suggest that attention mechanism type is a potentially deployment- relevant consideration, not merely an implementation detail.