Where Zero Coverage Stops Binding: Single-Shot Routing from 0.27B to 4.3B
Abstract
Routing among small language models is often motivated by the oracle gap: the accuracy a perfect per-query selector would gain over the best single model in a fixed pool. We ask what bounds that gain when every model is small. Across 22 instruction-tuned models from 0.27B to 4.3B parameters on four benchmarks, the binding constraint is zero coverage: items that no pool member answers correctly. In two-model pools near 0.3B, 43–63% of items receive no correct answer from any member in one attempt each; by 4B this falls to 5% on GSM8K but remains 22% on MATH-500. On free-response tasks the normalized oracle gap grows with scale under every gating regime we test, though its magnitude is sensitive to which bands are included. It grows because the uncovered set shrinks, not because models become more complementary: the share of items one model rescues for another stays roughly flat. We also show that near-chance members can produce large apparent oracle gaps on multiple-choice tasks, and give a difficulty-aware permutation test that detects this.