Auditing Selective Trust: A Pre-Deployment Benchmark and Mitigation for Trustworthy AI
Abstract
Language models increasingly answer questions using evidence they did not choose: a retrieved document, a user-provided claim, or a search snippet. One plausible but misleading signal can make an otherwise correct model endorse a wrong answer; yet a model that rejects all context is equally unsuitable because it cannot benefit from reliable evidence. We frame this deployment-relevant tension as selective trust. We introduce MIST, a human-annotated, controlled pre-deployment audit that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric that counts how often a misleading signal flips a clean-correct answer to wrong. Across the 23 models in our comprehensive auditing block, every model exhibits nonzero susceptibility, showing that standard accuracy alone is insufficient evidence for trusting a model with external signals. We then propose SCOPE, which mines clean-correct/misleading-wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Across both trained model families, SCOPE is the only evaluated mitigation whose point estimates do not reduce any of the three control accuracies relative to its base model; every alternative reduces at least one. It substantially lowers SC2W relative to each base, and the learned behavior transfers zero-shot to external benchmarks. For evidence-grounded systems used in AI-for-good settings, trustworthy deployment should be evaluated on selective trust, not resistance alone.