UniForm: Segment-Aware GRPO for Joint Punctuation Restoration and Inverse Text Normalization
Abstract
Punctuation Restoration and Inverse Text Normalization (ITN) are essential post-processing tasks for ASR systems. Traditionally, these are handled by separate models in a sequential pipeline, which suffers from error propagation and suboptimal latency. We propose UniForm, a unified framework that jointly solves both tasks using a single compact language model. UniForm is trained via a three-stage progressive paradigm: large-scale supervised fine-tuning on punctuation data, joint fine-tuning with ITN span masks, and reinforcement learning alignment. To enable effective reinforcement learning in this inherently low-entropy setting, we introduce Segment-Aware GRPO (SA-GRPO), which optimizes model performance through two synergistic components: a structured reward decomposition mechanism that provides dedicated signals for distinct segment types, and a token-level adaptive credit assignment mechanism that dynamically redistributes gradients toward difficult tokens based on sampling consensus and model entropy. This integrated approach effectively eliminates the need for manual sub-task weighting and ensures fine-grained optimization for both tasks. We train UniForm-0.8B on over 238 million multilingual samples. Extensive evaluations across diverse languages, domains, and code-switching scenarios demonstrate that UniForm-0.8B outperforms prior specialized models and larger general-purpose LMs on both tasks while maintaining real-time inference latency. Furthermore, our analysis reveals a strong synergistic effect, proving that joint training mutually benefits both punctuation and ITN performance.