Multilingual Speech Synthesis from Raw Bytes: Pronunciation, Language Identity and Code-Switching
Abstract
We build a multilingual speech synthesizer that connects three representations through a 65.7M-parameter trainable renderer: contextual UTF-8 bytes from a frozen text encoder, measured delivery coordinates, and low-rate tokens from a frozen speech codec. The system supports seven languages in six scripts, two target voices, delivery control without reference audio, and streaming audio. The frozen text encoder and codec add parameters beyond the trainable count. Segmented synthesis achieves mean per-language character error rate (CER) 0.0395 on 420 held-out sentences. Anchoring the delivery posterior to measured statistics and training the renderer to use individual coordinates gives four measurable pitch and energy controls. Matched Hindi experiments show that contextual native text reduces CER compared with eSpeak-derived phonetic input and static byte embeddings. Streaming produces the first audio in a median 153-160 ms after text encoding, with higher CER. Code-switching experiments show how added training data improves English-word recall, with different gains across languages. These results show how a compact renderer can use pretrained linguistic and acoustic representations while supporting explicit delivery control.