Answers Without the Model: Harness Failures and a Zero-Cost Noise Floor for Small-Scale LLM Evaluation
Satoshi Akiyama
Abstract
When an LLM evaluation harness is written by the same person who reads its output, its defects arrive disguised as findings. We report three such defects measured inside one continuously-logged evaluation programme (20 probe waves, 79,021 rows, one consumer GPU): a benchmark condition whose distractors reconstructed the gold answer, an item set whose key contradicted its own text, and a scoring vocabulary that counted non-answers as wrong answers. Each produced a clean, publishable-looking number. We then show that the same logs contain, for free, a calibration that would have caught two further retracted claims: repeated measurements of identical cells give an over-dispersion of $\phi = 1.01$ across eight waves, which converts to a minimum readable difference at any $n$. Applied retrospectively the threshold rejects both of our small-n retractions and passes the claim that survived re-measurement. Applied to all 46,279 cell pairs it finds only 17.3% readable. Finally, preparing this paper produced a fourth failure of the same family—eight of forty-two reported figures had drifted because the corpus was still being written—which we report as a case rather than silently correct.
Chat is not available.
Successful Page Load