What Does a Constant Value Profile Mean? An Identification Problem in Cross-Lingual Value Audits
Abstract
Value alignment is specified and audited in English, and a growing literature asks whether it transfers across languages by comparing a model's value profile across languages and treating small cross-language distance as stability. We ran that audit on three frontier models over 20 contested propositions in 13 languages and report a methodological result that matters more than the effect we set out to measure: the profile-distance statistic is not identified without a response-distribution diagnostic. One model returned the neutral midpoint on 99.3% of contested items; its cross-language distance is 0.009 Likert points, the smallest in our study, so any profile-distance audit would rank it the most cross-lingually stable system tested. Whether that reflects an absence of positions or a policy of not stating them is not identified by the statistic. We run the arm that separates them: with the midpoint removed and refusal offered explicitly, that system declines on 278 of 280 items while a dispersed system never declines. The constancy is therefore a policy, not an artefact of offering a neutral option. Reading the rating distribution directly on open weights shows the remaining question is answered per model rather than in general — one model puts 0.971 of its mass on the midpoint it selects, another only 0.483 with 0.352 on a competing option. A mean-based profile is degenerate in the opposite direction too, since a bimodal responder and a midpoint responder share a mean and therefore a distance. Among models that do produce dispersed ratings, distance differs by 0.186 Likert points, a factor of 1.8 (95% CI [+0.041, +0.336]), and the hypothesis that a minority language's values are pulled toward its dominant sibling is not separable from zero on the models whose responses disperse (+0.052, 95% CI [-0.031, +0.124]); pooling in the degenerate model attenuates that estimate by 27%, a demonstration of our own thesis. We document three ways such an audit silently yields meaningless numbers — degenerate response distributions, comprehension failure, and an automated translation reviewer repairing a control's truth value — with the checks that caught each, and release all responses and the deviation log.