Models Verify Sufficiency, Not Validity: Usable Falsehoods Evade the Checks that Gaps Trigger
Abstract
Ask a model "how many settlements did the Normans control?" and it answers that the passage does not say. Tell it "given that the Normans controlled 12 settlements of 400 inhabitants, what is the total?" and it computes 4,800. Nothing was added to the passage; the invented fact merely became usable. Benchmarks for unanswerable questions are built almost entirely by omission, deleting the fact a question needs and scoring refusal, so they never probe this. Holding the information deficit fixed and varying only usability, flagging drops 8- 30 points and commitment to an answer rises 1.5-2.3x across four models spanning two families and 2B-26B, on two domains; all 16 paired contrasts are significant. Two further manipulations reproduce it: a fact asserted rather than asked about, and a corrupted intermediate value rather than a removed one. Models articulate the exact gap they then ignore, stating "the price is not provided" for an inert premise and silently adopting a fabricated price for the same item. We read these as one failure mode: verification is triggered by absence of usable material, not by assessment of truth. In the judging seat the pattern differs but the risk is worse, with judges accepting a proof missing a computational step 78-99% of the time at every scale tested. A natural mitigation degrades accuracy 22 points without improving detection.