Freeze-Omni for Modern Standard Arabic: A Comparative Evaluation of Frozen-LLM and Cascaded Speech-to-Speech Architectures
Abstract
Speech-to-speech systems for Modern Standard Arabic (MSA) have been far less studied than their English and Mandarin counterparts. This paper introduces an MSA adaptation of Freeze-Omni, a frozen LLM speech-to-speech architecture additionally fine-tuned in both fixed-chunk and dynamic-chunk configurations on 2115 hours of MSA question-answer speech data synthesized using Microsoft Edge TTS. The model is compared against a traditional cascaded pipeline built on production Arabic speech components (Intella Voice for ASR and Zilla for TTS) with fully controlled experimentation using the Qwen2-7B-Instruct backbone across both architectures. Both models are evaluated on a held-out test set of question-answer pairs using automatic metrics (word/character error rate for speech understanding and generation; relevance accuracy and hallucination rate with LLM as judge analysis; and per-utterance latency), supplemented with human evaluation using naturalness, intelligibility, and prosody scores along with a comparative mean opinion score. Moreover, statistical significance testing throughout is conducted. The results showed that the end-to-end system achieves significantly lower word/character error rates and much lower response latency than a traditional cascaded system, while also achieving higher content relevance accuracy, a lower hallucination rate, and higher perceptual intelligibility and prosody. These results show that the frozen-LLM paradigm currently trades content accuracy and audio quality for sizable gains in response time and system efficiency, making it a promising low-latency option for MSA rather than an unconditional replacement for cascaded pipelines at current data scales.