Opt.Gear: Designing Hybrid Language Models for Device-Native Inference
Abstract
Deploying foundation models on devices requires more than reducing parameter count: autoregressive inference must also respect constrained memory bandwidth, persistent state, and the operator support of heterogeneous CPU, GPU, and NPU runtimes. We provide a systems-aware characterization of Model H, a device-native language-model family that replaces many local-attention layers with a Convkv-Gated Mixer. The mixer uses causal depthwise convolution and element-wise gating, maintaining a fixed-size local state rather than a sliding-window key-value cache; a small number of grouped-query attention layers retain global information routing. At the 1B scale, Model H is trained on 0.5T Korean-English tokens without teacher-model distillation and supports a 64K context length. We document its task-capability profile, post-training quantization behavior, and measured deployment throughput. Under W4A16 deployment, Model H-1B reaches 7,042 prefill tokens/s and 86 decode tokens/s on a Snapdragon Gen 5 NPU through Qualcomm AI Runtime, while retaining competitive English and Korean benchmark performance. Measurements on Qualcomm and Apple stacks show that the reported design can map efficiently to multiple device-specific runtimes, although absolute behavior is runtime dependent. The evidence characterizes this reported configuration and does not isolate architecture-only causal effects.