When Language Rewrites the Answer Source: Typed Arbitration and Subadditive Composition in Foundation Models
Abstract
A language model can answer the same false-belief scenario correctly for different reasons: it may follow the queried agent, copy reality, track another agent, or rely on a distractor. Accuracy alone cannot distinguish these possibilities. We study this problem as typed source arbitration: how linguistic form changes which information source governs the answer while the underlying event and candidate answers remain fixed. By crossing matched descriptions with controlled conflicts among these sources, we show that linguistic effects are conditional on the competing sources rather than stable, context-free answer biases. We then ask whether multiple linguistic constraints combine independently. Freezing the effects of individual constructions before evaluating unseen compounds reveals a systematic departure from additive prediction: in the focal instruction-tuned checkpoint, the joint effect is smaller than the additive prediction on independently generated scenarios. The same directional pattern does not replicate across other checkpoints, and a rescaling test on Qwen3-8B rules out confidence scale alone as the explanation. These findings identify a checkpoint-contingent pattern in how linguistic descriptions combine. More importantly, they distinguish a model that reaches the correct answer from one whose answer is governed by the intended belief source, preventing accuracy-only evaluation from mistaking reality-following or distractor-driven success for belief tracking.