ixi-sLM: On-Device sLM for AI Call Agent Service
Han-Sang Lee ⋅ Kyuho Lee ⋅ Jin K Park ⋅ Minjae Kim ⋅ Seungheon Hyeon ⋅ KIDUK KWON ⋅ Juneyoung Park ⋅ Seongbae Lee ⋅ Yuri Hong ⋅ Jinwoo Lee ⋅ YOUNGWOOK KWON ⋅ Seongwan Kim ⋅ Jaeho Lee
Abstract
Running an AI call agent on-device makes it private, network-independent, and free of per-query cloud cost, but shipping one as a real-world service imposes constraints a cloud deployment never faces: $\sim$1 GiB memory budget, in-call latency and thermal limits, static-integer NPU execution, Korean STT input far from sLM pretraining, and cloud-level quality expectations. We bring a 1.2B Korean sLM to mobile NPUs through on-device modeling: domain-adaptive modeling trims the vocabulary by call frequency to shrink the embedding and LM head that take nearly half the model under INT4 transformer weights and refits the tokenizer to call speech (18\% fewer tokens on call traffic), while task-aware quantization calibrates on real calls and assigns per-module mixed precision by activation sensitivity. A single domain-adapted model runs two production services, AI call summary and AI call answering, at 94\% of a state-of-the-art cloud LLM's task-weighted quality; our on-device modeling carries that quality onto the NPU essentially without loss, in an 803\,MiB model (28\% of FP16) at 23\% of the CPU's per-call energy. The same recipe deploys across Qualcomm and Apple NPUs and serves several million subscribers in production.
Chat is not available.
Successful Page Load