Epistemic Arbitration in Language Models
Abstract
Language models can recognize that an answer is wrong and still choose to use it up to 99.4% of the time. Understanding when models accept or reject information from tools, retrievers, other agents, and users is therefore a distinct problem from understanding whether they can evaluate that information correctly. We call this decision epistemic arbitration. To explore this phenomenon, we created over 10 million trials on 12 language models across the Qwen, Gemma, Mistral, and Llama families. We find that these properties manifest across compositional arithmetic, natural-language quantitative reasoning, linear algebra, propositional constraints, physics, and molecular biology. Across all domains, models systematically favor answers that align with their latent priors. We also find that the same answer can receive very different weight depending on how its source is described, and that the same advice can significantly help or harm depending on the model's competence. These arbitration policies vary substantially across models and domains. In propositional constraint reasoning, models still adopted candidates they had correctly rejected in isolation 93–100% of the time. In mechanistic case studies, causal interventions reveal a staged process by which external answers are integrated into the model's computation, while verification evidence can be represented and even verbalized without governing the final answer. Our results suggest that reliability with external information is not just a matter of capability, tool accuracy, or verification ability, but instead requires understanding how external information is trusted and used.