Knowledge Distillation from a Raw-Signal Foundation Model to a Feature-Based On-Device Sleep Stager
Lucas F Buzuti ⋅ Daniel I Futata ⋅ Jhonatas Conceição ⋅ Felipe Chaud Pinheiro ⋅ Jean P de Matos ⋅ Byeong-ho Lee ⋅ Seungman Yang ⋅ William Camilo Ariza-Zambrano
Abstract
A foundation model pretrained on raw wearable biosignals cannot run on the wearable. It reads photoplethysmography and accelerometry at full rate and carries hundreds of millions of parameters, while the device computes only cheap derived features and has room for a model orders of magnitude smaller. We distill across that gap, from a 181M-parameter teacher into the 0.69M-parameter classifier already deployed on the wearable, whose architecture and runtime are fixed by the product. The teacher reads spectrograms of the raw signal and the student inter-beat intervals and accelerometry, so nothing maps one input back to the other and the teacher's posterior is the only channel between them. On four-class sleep staging the student reaches 72.11 macro F1 and Cohen's $\kappa$ of 0.66, against 70.04 and 0.63 for the traditional supervised model it replaces and 72.99 and 0.67 for the fine-tuned teacher, recovering the teacher's quality at 0.4% of its parameters. Balanced accuracy is unchanged, 78.63 against 78.48. One such student, run on the real-world pipeline and scored against polysomnography, improves on the traditional supervised model by 2.19 macro F1 over 407 watch recordings, and is level with it over 44 recordings from a ring it was never trained for.
Chat is not available.
Successful Page Load