Position Bias Masquerades as Moral Steering in Multilingual Forced-Choice Evaluation
Sharvin Goyal ⋅ Jeremy Kalfus ⋅ Ajay Agarwal
Abstract
Prior work steers moral frameworks across languages and reports that lower-resource languages are harder to steer. We ask whether that reading survives controls for instruction following, answer validity, and option-order sensitivity. On the MultiTP trolley-problem corpus we run seven model-mode cells, each with 19,320 generated responses across six languages and seven conditions. Under the standard forced-letter format, baseline same-pole agreement under a label swap falls below the 0.5 reference in six of seven cells, and any reported consistency rate decomposes into complementary same-letter and same-pole readings. In a baseline-only ablation on two models, requesting the chosen option text instead of its letter raises paired same-pole agreement by $+0.27$ to $+0.62$ and recovers parseable Zulu pairs from a model that produced none under forced letters. Restricting to order-stable pairs retains 5.5 to 43.0% of scheduled pairs, moves resource-gap estimates in both directions, and leaves interval-exclusion counts that depend on the resampling unit (8 of 28 with independent records, 11 of 28 with dilemma clusters), which we report as exploratory. A pre-specified safety comparison on one cell gives a Zulu-minus-English attack-success gap of $+0.083$, below its 0.15 threshold. Multilingual steering scores should be reported with response validity, semantic order sensitivity, selection coverage, and the scope of their controls.
Chat is not available.
Successful Page Load