GoodEnough: A Pre-Registered Non-Inferiority Study of a Quantized 1.7B Model on Consumer CPU Against a Hosted 70B Model
Abstract
Cost-aware routing between a small model and a large one is a standard proposal, but the routing literature mostly compares hosted models against other hosted models, and where it does not, the local tier is unquantized or on accelerator hardware. We ask instead where a quantized 1.7B model running on a consumer laptop CPU can replace a hosted 70B model, and we fix the answer’s format in advance: a public pre-registration commits the margin, the two primary benchmark slices, the statistical test, and an explicit falsification condition before any data is collected. On none of eight MMLU slices does the local model establish non-inferiority at a 10-point margin, and both primary slices fall below it, so the pre-registered hypothesis is rejected. The more useful result is structural. On a held-out split the local model answered correctly where the hosted model failed on 5 of 140 items, bounding the accuracy a selector over those two answers could recover over always calling the hosted model at 3.6 points, two-sided 95% interval [0.7, 7.1]. A parse-failure cascade escalated on 4 of 140 items, because the local model fails by returning well-formed wrong answers rather than nothing. Total measured cost was $0.26.