What Does Language Alter Medical Vision-Language Model Decisions Across Domains? A Seven-Language Chest X-Ray Study
Ayoub Louaye Bouaziz ⋅ Douaa M Yagoub ⋅ Iliam Souami ⋅ zein miled ⋅ Ilyes sais ⋅ Nadjiba Rahal ⋅ Basma Halil ⋅ Chemam soundous ⋅ REZZOUG AICHA ⋅ Lokmane Chebouba
Abstract
Medical vision-language models (VLMs) are increasingly evaluated as image-conditioned clinical assistants, yet prompt language is often treated as a neutral interface and acquisition domain as a separate robustness problem. We evaluate these factors jointly on a chest X-ray benchmark spanning 11 data sources, 35 RadLex concepts, seven languages, and ten VLMs. Because model coverage is uneven across the full archive, we predefine a complete-case reporting rule: cross-model claims use only dataset--pathology cells containing all ten models in all seven languages. This yields a 12-cell matched core across four domains and 69,388 prediction attempts. On the stricter binary-complete core, mean balanced accuracy falls from 0.668 in English to 0.521 in Arabic, 0.549 in Hindi, and 0.595 in Chinese. On the matched accuracy core, the corresponding changes relative to English are $-18.2$, $-15.9$, and $-15.6$ percentage points. Language sensitivity is highly model-dependent: the across-language accuracy range varies from 0.0005 for Qwen2.5-VL-7B to 0.577 for Lingshu-7B. Crucially, Qwen2.5-VL-7B combines apparent language stability with balanced accuracy near chance and F1 of 0.016, showing that consistency alone is not evidence of robust medical reasoning. Model rankings also change across languages (mean Kendall $\tau=0.444$ on complete ten-model domain comparisons). These results show that language and acquisition domain must be crossed explicitly when evaluating medical VLM robustness.
Chat is not available.
Successful Page Load