From Breath to Command: A Personalized and Explainable Pipeline for Augmentative Communication
Abstract
Individuals with severe motor impairments, including amyotrophic lateral sclerosis, locked in syndrome, and high level spinal cord injury, often retain voluntary respiratory control, making breath a viable channel for assistive communication. However, existing breath based augmentative and alternative communication (AAC) systems rely on fixed pressure thresholds, use population level classifiers that do not adapt to individual users, and offer no interpretability for clinical validation. We present an end to end breath gesture recognition pipeline, shown in Figure 1, that addresses these limitations. Raw microphone audio is segmented and converted into log Mel spectrogram features, which a Multi Scale Temporal Convolutional Network with four parallel dilated branches (dilation rates 1, 2, 4, and 8) encodes into a shared embedding, capturing both short onset transients and sustained exhalation patterns within the same architecture. Prototypical Networks then personalize the classifier from a small per user calibration set without retraining the global model, while a dual pathway Grad CAM module explains both the global and personalized decisions. Evaluated under strict subject wise splits on the Coswara respiratory sound dataset (2,542 subjects), the global model achieves macro F1 of 0.726 on 382 unseen subjects, outperforming CNN (0.724) and CNN BiLSTM (0.696) baselines trained on the same split. Cross subject personalization with 10 calibration samples per class reaches macro F1 of 0.730 without retraining, and an augmentation based within subject simulation establishes an upper bound of 0.984. Multi scale branching contributes a consistent 0.013 macro F1 improvement over a single scale baseline, and inference latency is 2.6 ms, well within real time requirements for assistive interaction. Grad CAM overlays further confirm that short and long breath gestures activate acoustically distinct, physically plausible spectrogram regions, supporting the reliability of the model's predictions. Prior breath based AAC systems have relied on fixed thresholds and population level models without adaptation or explanation, and respiratory audio classification and few shot personalization have largely been studied separately, with explainability for breath based interaction specifically remaining underexplored. To our knowledge, this is the first system to combine multi scale temporal modeling, few shot personalization, and dual pathway explainability in one deployed pipeline, with clinical validation on motor impaired participants identified as the necessary next step.