Truthfulness Is Not Enough: Prompt-Conditioned Evidence Selection in LLMs
Abstract
An AI assistant can say nothing false and still leave a user wrong. Responses can be equally truthful but have very different epistemic consequences: some expose a user's mistaken model while others leave it intact. Standard truthfulness evaluations collapse this decision problem: once many truthful responses are available, factuality metrics cannot tell us which truth a model should select for a particular user. We introduce EpiSelect, an evaluation of the epistemic policy governing that choice. The model knows the true rule, observes another agent's judgments, and generates new evidence under Corrective, Default, or Accommodation objectives. A deterministic evaluator separately scores user-rule inference, truthfulness, and corrective coverage over the user's surviving hypothesis space. Across Boolean and relational rule domains, three frontier models (Claude Opus 5, Gemini 3.1 Pro, and GPT-5.6 Sol) achieve perfect Boolean user-rule inference and near-perfect truthfulness, yet explicit objectives move their selections across nearly the full attainable range: corrective prompts approach the most diagnostic evidence available, while accommodation prompts shift sharply toward evidence under which the user's apparent rule remains viable. Default behavior falls between these endpoints on average, though it reaches the corrective ceiling in some worlds. Because inference is solved, this span reflects the selection policy rather than a capability limit. By contrast, an 11-model open-weight sweep fails to reliably realize either endpoint, separating capability from policy control. These results identify evidence selection as a distinct, prompt-controllable layer that standard truthfulness evaluation leaves unmeasured, and that shapes whether an assistant surfaces a user's mistake when the user does not know to ask.