Testing by betting when the bets are real: an e-value audit of 84 pre-registered hypotheses about a deployed forecaster
Prigodskii Roman
Abstract
E-values are motivated by a gambler betting against a null hypothesis. We report an audit where the gambler is not a metaphor: the null is a bookmaker's posted price, the stake is a stake, and the wealth process *is* the e-process. The subject is a deployed UFC bout-outcome forecaster and a pre-registered search for pockets where it beats the closing line: 84 hypotheses, 1,787 priced bouts, first analysed with paired log-loss and Benjamini–Hochberg. Re-asked as bets, three things change. (i) The 84 slices are overlapping subsets of one pool, so BH's guarantee is not established; its one favourable discovery fails under both Benjamini–Yekutieli and e-BH. (ii) A segment the fixed-$n$ analysis reported as *failing to replicate* ($p=0.30$) is, on the wealth scale, still growing in the replication window and 100 bouts short of $1/\alpha$ at fair odds, 253 at the book's: optional continuation converts a refutation into a schedule. (iii) A post-hoc threshold rule, paid for by a mixture martingale, survives the search it actually ran ($e=31.3$) and dies twice: when the direction it could also have searched is priced in ($15.6$), and when the bookmaker's margin is ($8.5$). We report the ladder, not the winner. The three failure modes we price here — overlapping evaluation slices, a follow-up window read as a refutation, a threshold chosen after the sweep — are ordinary habits of empirical ML, and the market only makes their cost legible.
Chat is not available.
Successful Page Load