Resource Inversion: Cross-Lingual Fidelity and Verified Deferral for Oncology Patient Communication in Low-Resource Languages
Abstract
The languages spoken by the populations with the highest burden of liver and bile-duct cancer receive the least language-technology support. Liver-cancer incidence in Mongolia, Laos and Cambodia runs three to thirteen times the US rate, yet Lao Wikipedia is three orders of magnitude smaller than English Wikipedia. We call this mismatch resource inversion and study what it means for a task cancer centres are beginning to automate, namely drafting plain-language after-visit summaries from English oncology notes in the patient's own language, under the deployment constraint of small open-weight models on on-premise hardware. Our benchmark decomposes 12 hepatobiliary notes into 72 typed clinical facts with gold renderings and trial-traceable doses, spans 10 burden-selected languages, and is complemented by a real-text set drawn from published case narratives. Because native-speaker raters are unavailable for most of these languages, fidelity is measured by back-translation entailment under an engine independent of each system, and the metric itself is validated with perturbation tests and per-language ceilings computed from FLORES human references. Fact-by-fact scaffolded drafting brings higher-resource languages to 80% of the metric ceiling. The lowest-resource languages, in contrast, cannot be written directly by 1.5-1.7B models at all; they reach 8% of ceiling amid Thai-for-Lao substitution, on-script degeneration and invented drug names. A workable division of labour remains available even there. The small model drafts plain English, a dedicated translation model renders the target language at 76-88% of ceiling, and a verification gate defers any fact it cannot confirm to a human interpreter. In higher-resource languages the gate removes 80% of unfaithful facts at a 47% deferral rate. In the floor languages, where 93% of drafted facts are unfaithful, it defers nearly everything, so the system fails toward human review rather than toward confident error.