Country Is Not Enough: LLMs Recognize Missing Cultural Context but Still Generalize
Abstract
Large language models often answer culturally situated questions at country scope even when a reliable answer depends on subnational, community, or situational context. We study cultural answerability: whether the scope supplied by a user supports a direct answer without promoting one situated practice to a national norm. From region-specific INDICA evidence, we construct a held-out bilingual English-Hindi evaluation with 50 culturally stable and 50 culturally dependent questions. Two multilingual LLMs are tested at three stages: explicit recognition of missing context, natural response policy, and paired regional counterfactual grounding. Models flag 89.5% of dependent conditions but correctly leave only 28.0% of stable conditions answerable. Under natural assistance, 30.5% of dependent responses make an unscoped country-level commitment. Even among conditions explicitly recognised as dependent, 31.3% still make that commitment. Supplying a region does not solve the problem: strict regional grounding succeeds in only 13.5% of paired conditions. Cultural reliability is therefore not only a knowledge problem; it requires calibrated recognition, scope-aware response policy, and grounded adaptation.