Even Irrelevant Medical Context Weakens Refusal: Cross-Lingual Safety Failures in Medical RAG
Abstract
Retrieval-augmented generation (RAG) is widely used in medical assistant systems powered by small language models, but its safety effects in cross-lingual use remain unclear. We study 12,000 outputs from four small instruction-tuned language models on 500 MedSafetyBench prompts and their Hindi translations under three conditions: no added context, query-relevant retrieval, and medically framed but query-irrelevant retrieval from a shared English MedRAG corpus. Using an LLM judge calibrated against human annotations, we find that harmful compliance is higher for Hindi than for English when pooled across conditions (44.8% vs. 35.4%), and rises from 30.8% with no retrieval to 40.7% with medically framed irrelevant retrieval and 48.9\% with relevant retrieval. A benign non-medical control does not reproduce this pattern, suggesting that medical framing, rather than context length alone, drives the degradation. Medical RAG systems built on small language models should therefore be evaluated for retrieval-induced safety failures across languages, not just retrieval accuracy.