MITRA-Swarm: Injecting Explicit Linguistic Structure into Foundation Models for Classical Machine Translation
Kayshav Bhardwaj ⋅ Pragun Seth ⋅ Kush Bhardwaj ⋅ Vijay Nagasamy ⋅ Kurt Keutzer
Abstract
Classical Sanskrit poses three linguistic challenges monolithic foundation models do not resolve in one generation call: morphophonemic fusion (Sandhi), dense nominal compounding (Samāsa), and discourse-level zero-anaphora forced by poetic meter. We present MITRA-Swarm, a multi-agent architecture injecting explicit linguistic structure into an otherwise monolithic pipeline: a Sandhi-splitting agent, a parallel array of generalist (gemma-4-26B-A4B-it) and domain-specialized (gemma-2-mitra-it) translators, and a discourse-aware Ellipsis Restorer that consults corpus-derived register tags to avoid inventing active-voice subjects in impersonal legal and philosophical registers. Quality is estimated by two strictly separated judges: a routing judge ($J_{\text{route}}$, same model family as the generators) driving an internal critique loop, and an independent reporting judge ($J_{\text{report}} = \text{gpt-4o-mini}$, no shared vendor lineage) scoring every output exactly once. On a stratified benchmark of 1,000 verses across five registers (23,990 scored verdicts, 12-condition pre-registered ablation), the full swarm reaches $85.90$ MQM against a monolithic baseline of $85.27$ ($\Delta = +0.63$, $p_{\text{adj}} = 0.0018$, 95% CI $[+0.30, +0.97]$), and isolated discourse restoration alone accounts for $+0.44$ MQM ($p_{\text{adj}} = 0.0288$), outperforming a call-count-matched self-refinement baseline by $+0.51$ MQM ($p_{\text{adj}} = 0.0288$). Severe per-verse failures (MQM $< 80$) fall from 9.0% to 7.6%, at roughly $9\times$ the per-verse inference cost of the baseline. Beyond these headline effects, we report two findings about evaluation itself: $J_{\text{route}}$ scores its own family's outputs 2–7 MQM higher than $J_{\text{report}}$ on every one of the 12 conditions, empirically motivating the two-judge split; and a preflight screen of a 192,151-entry retrieval index against all 12 evaluation works verifies 0 reference leaks across 269 runtime checks. We report these as small effects on a metric that saturates near 85–86 across every condition.
Chat is not available.
Successful Page Load