Multilingual Rule Portability Fractures by Required Response Act: Which Directives Survive Translation, and Which Stop Binding
Anu Adesina ⋅ Aviram Berg ⋅ Ayesha Imran ⋅ Janhavi Khindkar ⋅ Maria Lomaeva ⋅ Ifeoma Okoh
Abstract
Multilingual rule following is often treated as a resource problem: as language-resource availability falls, rules are expected to weaken broadly. We show that this account is incomplete. We evaluate matched directives across ten languages spanning public NLP-resource availability, seven rule categories, and two 8B instruction-tuned models. In the two comparatively lower-resource conditions, rules satisfiable by withholding content average 96.3\% HELD, whereas rules requiring a constructed referral or reasoned refusal average 7.9\% in the same runs. The observed lower-resource failure is therefore selective by required response act, not uniform across a language. Across 140 category--language--model cells, rule category accounts for 47.2\% of descriptive variation and category $\times$ language for 23.2\%, compared with 21.5\% for language; after removing a deliberately difficult inversion control, category $\times$ language is largest (38.3\%). Internal analyses reveal a complementary dissociation. Coarse obligation and active-rule-state contrasts remain linearly decodable in weak-binding conditions, but the finer \emph{must}--\emph{may} direction transfers poorly across unseen grammatical framings where adherence is weakest and covaries with HELD in both models ($\rho=.73$ and $.74$). This argues against a pure missing-information account and is consistent with a form-sensitive binding or execution gap; probes and correlations do not establish causal use. Because only two languages instantiate the lower-resource end and response-act classes differ in task and rubric complexity, we do not claim a universal resource effect. We contribute a factorial multilingual benchmark and an evaluation framework that separates three questions normally collapsed by pass rates: whether rule information is detectable, whether the model engages with the constraint, and whether the rule controls the required response. This shifts multilingual certification from a language average to a rule--language--model combination and motivates testing form invariance rather than assuming that English success transfers.
Chat is not available.
Successful Page Load