The Split Decides the Score: Partition Instability in Cross-Region Crop Yield Benchmarks
Abstract
Leave-one-region-out evaluation is a common test of whether crop yield models generalize geographically, and recent audits report severe out-of-distribution failure. We show that on one such benchmark, several headline conclusions depend on an under-examined design choice: how the map is split into regions. Evaluating seven methods on 8,691 county-years of U.S. corn under an identical leakage-controlled protocol, we vary the fold definition across a state-level approximation of USDA Farm Resource Regions (one we initially adopted ourselves), the official county-level regions, and climate-derived clusters. The highest-scoring method differs across all three, zero-shot difficulty ranges from −5.48 to −0.43 macro R², and the state-level approximation misassigns 44.5% of county-years, omits three of nine official regions, and has a most influential fold that is 30 of 31 counties mislabeled. Transfer performance is associated with the held-out region's own predictability (r=+0.74 across 22 non-independent cells, partly mechanical), while neither covariate- nor label-distance shows leverage-robust association. The one cell that makes distance look predictive is itself an artifact of the incorrect partition. Climate conditioning, unlike meta-learning, has the higher mean in all six feature×partition settings. We release a support-matched harness, the county crosswalk, and reporting recommendations.