Failure-Search Rankings Depend on the Requested Repertoire
Abstract
Which search method finds the most useful failures depends on what its output must contain. We study frozen Chronos-Bolt small and base models under 144 controlled temporal perturbations, holding logical query budgets, initialization and output extraction constant across search methods. An initial 384-search experiment on M4 weekly and synthetic data finds that archive mutation improves eight-cell failure repertoires at 64 queries; a 192-search follow-up also beats score-blind balanced sampling. We then prospectively test output-objective dependence on 128 new seasonal series, with 32 search and 96 transfer series and 240 new searches. Changing only the final suite's minimum descriptor quota from eight cells to one reverses Archive versus single-elite rankings in both model sizes. The paired quota interactions are 0.1083 and 0.1251, with both six-comparison-adjusted intervals above zero. A simple policy sampling parents from the current objective-constrained suite improves unconstrained Top-8 transfer scores over Archive by 0.1146/0.1173, but has no established advantage for eight-cell repertoires. Every suite is sealed before transfer forecasts are generated. The evidence supports matching search to the requested failure set, rather than treating a repertoire score as general optimizer superiority.