Continuous Ethiopian Sign Language Recognition: A Multi-Modal Study of Video and Skeleton-Based End-to-End Models with Cross-Lingual Transfer
Abstract
Sign language is a primary means of communication for Deaf communities, yet the limited availability of automatic recognition systems creates significant barriers to communication and access to information. In Ethiopian Sign Language (EthSL), existing research has largely focused on isolated signs, while continuous sentence- level recognition remains underexplored because of the temporal complexity of signing and substantial variation across signers. To address this gap, we introduce a multi-signer dataset for continuous EthSL recognition comprising 2,220 video sequences of 50 unique sentences performed by 22 signers, with each sentence performed twice. We investigate two complementary end-to-end approaches for continuous recognition. The first is a video-based architecture that combines a ResNet-18 backbone for spatial feature extraction, temporal convolution and frame- correlation modeling for motion representation, Bidirectional Long Short-Term Memory (BiLSTM) for long-range temporal dependencies, and Connectionist Temporal Classification (CTC) for alignment-free sequence decoding. We further employ attention-based temporal pooling and controlled background and clothing augmentation to improve robustness to signer variation. The second is a lightweight skeleton-based architecture that uses MediaPipe hand and body keypoints with a spatio-temporal graph convolutional network (ST-GCN), BiLSTM, and CTC. To address the absence of pretrained African sign-language models, we additionally evaluate cross-lingual transfer using a frozen hand encoder pretrained on Saudi Sign Language. Under signer-independent and unseen-sentence evaluation settings, the video-based model achieves WERs of 8.82% and 58.0%, respectively. For the skeleton-based approach, the baseline ST-GCN + BiLSTM + CTC model (832K parameters) achieves 18.5% WER on signer-independent evaluation and 79.4% WER on unseen sentences. After integrating the frozen Saudi Sign Language hand encoder (1.9M parameters), performance improves substantially to 8.2% WER on signer-independent evaluation and 60.4% WER on unseen sentences, corresponding to absolute improvements of 10.3 percentage points and 19.0 percentage points, respectively. These results demonstrate that both visual and skeletal representations can support continuous EthSL recognition, while cross-lingual transfer provides a particularly effective strategy for improving recognition in low-resource sign languages. The study establishes a benchmark for continuous EthSL recognition and provides a foundation for developing efficient and accessible sign language technologies for automated translation and assistive communication.