Encoding Articulatory Features in SSL Models for Hindi Dysarthric Speech
Abstract
Self-supervised speech models (SSL) are increasingly used to assess dysarthria, but it remains unclear which aspects of articulation their frozen representations preserve once speech becomes disordered, and whether findings drawn almost entirely from English hold in a language with different phonology. We probe three SSL models (WavLM, XLS-R, and the Hindi-pretrained IndicWav2Vec) on isolated Hindi vowels and consonants from 13 healthy and 14 dysarthric speakers, decomposing each consonant into place, manner, voicing, and aspiration rather than treating it as a single label. Across all three models, voicing survives dysarthria far better than place or manner (macro-F1 drop of 0.12–0.15 vs. 0.34–0.44); once corrected for chance, however, place and manner still carry more usable signal than voicing (up to 0.42 vs. 0.26–0.32 above chance), with aspiration retaining the least. We also show, for the first time, that two contrasts unique to Hindi and absent from English, retroflex vs. dental place, and aspirated vs. unaspirated consonants, measurably collapse under dysarthria (significant in 5 of 6 model × contrast tests). The Hindi-pretrained model is associated with sharper representations of these contrasts, though with one model per pretraining condition we cannot isolate this from confounds in architecture and training scale. SSL representations thus encode a clinically recognizable signature of articulatory breakdown, with pretraining language shaping its strength but not its overall shape, arguing for testing speech-pathology tools beyond English-centric assumptions.