Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and Large Language Models
Abstract
Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three encoder-side diagnostics to characterize how training paradigms shape uncertainty reduction across the speech encoder-LLM interface. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a capability-boundary-aware multi-stage training strategy that targets parameter efficiency and robustness to hallucinations. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on various benchmarks show that our method achieves competitive performance against state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.