Position: The Language of Thought Is a Design Choice, Not a Compliance Problem
Abstract
One norm is hardening across multilingual reasoning, that a model answering a question posed in a non-English language should write its chain of thought (the trace) in that language too. Language-consistency rewards, \textit{think natively} objectives, and evaluations that score trace-language compliance all enforce it. We argue that the norm commits a category error by fusing two functions of a reasoning trace. As \emph{computation} the trace should choose its medium for correctness, and as \emph{communication} it should choose its medium for the audience. It resolves that conflict by taxing the first to subsidize the second, precisely for the languages that can least afford it. Evidence for that trade-off is already scattered across prior work that no one has read together, so the field lacks the argument rather than the data. We make that argument and ground it in 25{,}200 generations over two 8-billion-parameter models, four thinking-language policies, seven languages and two tasks, which return one result we have not seen reported elsewhere. Asking the stronger model to think in the target language barely changes its trace. Of the six non-English languages we test, only German is an exception. Native thinking nonetheless remains Pareto-dominated on accuracy against token cost in every typologically distant configuration. A mixed policy that reasons in English, keeps the question's entities and quantities verbatim, and answers natively delivers most of what the norm wants more cheaply. We then propose a decision framework for choosing the language of thought by task type, typological distance and audience, and a thinking-language policy card that asks releases to disclose and justify their choice rather than default to consistency.