What Claims Do LLM Benchmark Scores Support?
Abstract
LLM benchmark scores are often read as claims about factual accuracy, evidence use, instruction following, or logical reasoning. We study this score-to-claim conversion as inference over a completed evaluation record. Given benchmark items, prompts, outputs, scores, and native fields, which capability sentence is supported, which stronger sentence remains outside the evidence, and would that boundary change model choice? We formalize a report-use procedure, \emph{score-to-claim inference}: hold source units fixed, compare each claim-induced condition to benchmark-native controls or same-source invariance checks, and apply a declared finite-family support rule. In a fixed set of five open-weight model families, completed records support narrower sentences than endpoint readings alone suggest: supplied-context answer recovery remains separate from support attribution; local supplied-evidence verdict signals leave open-web provenance unmeasured; and SATBench exposes label-surface instability for SAT/UNSAT prediction. The output is a support boundary, a blocked strengthening, and the next measurement or model-selection sensitivity attached to the reported result.