When More Search Can Worsen Deployment: Best-of-N Selection on an Incomplete Proxy
Dev D Goyal
Abstract
A statistically well-estimated selection score can still be incomplete: it may omit part of the deployed objective. Wider best-of-N search then changes the winner distribution and can expose candidates with larger omitted terms. We show that proxy-only information cannot identify an unrestricted candidate-specific deployment term. For uniform sampling without replacement from a realized finite pool, an exact rank-exposure identity separates known reweighting from the empirical outcomes attached to ranks. In a trading case study, the capacity-sensitive search-induced gap is positive in all 16 primary pool-by-capital cells. For reference pool B at 500M, its 95\% joint selection--evaluation block-bootstrap interval $[-0.0219, +0.1209]$ includes zero. A post hoc net-score rule improves mean deployed Sharpe by 0.430 relative to gross-score selection on the tested grid; a paired block bootstrap gives $[+0.192, +0.723]$ for that gain, although the resulting mean remains $-0.071$. The mechanism and remedy are conditional on the realized paths, candidate pools, survivor-filtered universes, calibrations, and daily simulator, not a population law.
Chat is not available.
Successful Page Load