Consistent, Confident, and Wrong
Abstract
In 2003 a bridge across the Rhine, built toward the middle from both banks, failed to meet by 54 cm. Each survey was internally precise to the millimetre; Germany referenced sea level at the North Sea, Switzerland at the Mediterranean. The instructive part is not the sign error that doubled the gap: no amount of additional surveying on either bank could have caught it, because every measurement was a difference between two points on the same bank and the datum cancels out of every difference. We argue machine learning has datum problems and measure four, the last with an instrument borrowed from a field that met the same problem in the 1940s–50s and solved it. The sharpest consequence: ensembling across seeds returns a more reassuring number precisely as the estimate gets worse. In a synthetic world where the answer is known by construction, five independently seeded reward models reach pairwise accuracy 0.979 and cross-seed agreement 0.904 while their correlation with the quantity they exist to estimate is 0.041; models differing 16× in width agree just as much. Static token selection has an exact closed form for its bias that we confirm to 10⁻¹⁶, and a frozen mask sends held-out loss to infinity. Consensus corpora fail below a measured coverage threshold. Finally we rebuild the numerical analyst's instrument — backward error, how far the problem is from one your answer solves exactly — calibrate it against a backward-stable anchor and a random-number null, and read it on seven systems spanning 0–72% accuracy. Treating η = ∞ as censoring rather than as a missing value — a correction to our own submitted analysis — the most accurate system separates from the null under every test, its median wrong answer nearer the stated problem than the anchor is, while the accuracy-ordered split we previously reported does not survive. In each case the objective is a consistent estimator of something other than the decision-relevant quantity, so the error is bias rather than variance; and the standard diagnostics are functionals of that same objective, so they cannot see it. We report one experiment of our own that failed, as it ran.