Continual On-Device Personalization of Small Language Models from Hindsight
Abstract
Small language models deployed on phones must adapt as users' preferences evolve, yet mobile interactions rarely provide direct training labels. This challenge is general: a model often acts from current context, while later user activity reveals whether the decision was useful. We turn such hindsight into continual on-device learning through self-distillation. A frozen copy of the same model revisits each decision with later evidence and supplies a soft target for the policy that acted from the original context. The system balances exploration and exploitation: uncertain decisions test plausible alternatives to obtain feedback, confident decisions follow the current policy, and balanced replay combines each new lesson with earlier ones. On a sequential personalization task, the method improves accuracy and reduces cumulative regret relative to in-context learning (ICL), retrieval-augmented prompting (RAG), reward-based learning, and rejection fine-tuning. We further validate the learning loop on a physical device, where observed user activity produces model updates that persist after the app restarts.