Minibatch Starvation: A Silent, Provable Failure Mode in Group DRO’s Reference Implementations
Aditya Sanjeev
Abstract
Group Distributionally Robust Optimization (Group DRO) is meant to protect underperforming groups by upweighting them, but it sometimes does the opposite: it drives a minority group's adaptive weight to zero and collapses that group's accuracy, often to 0\%, while the aggregate metric barely moves, so a standard evaluation can certify a model that has silently failed the group it was built to protect. We identify the cause: the standard minibatch estimator *zero-fills*, assigning loss 0 to any group absent from the current batch, and the exponentiated-weight update then drains that group's share through the normalizer. We prove this estimator artifact is sufficient for collapse by construction (holding all group losses exactly equal, where the true update provably preserves uniform weights) and derive a closed-form collapse-time law, $T_{collapse} = \Delta^\circ/(\eta\bar\ell(1-f)^B)$ with $f = p/(1+R)$, that reproduces the measured phenomenology from one calibrated scale and explains an apparent dataset-size effect as a unit artifact of measuring in epochs rather than gradient steps. We audit three widely used reference codebases (the official Sagawa et al. implementation, WILDS, and SubpopBench) and prove from source, line by line, that all three carry the mechanism on their default training path. A kill test confirms it in vivo: renormalizing the weight update over only the batch-present groups suppresses collapse by more than an order of magnitude; the naive fix, skipping the absent group's own update, provably does nothing, because the artifact acts through the shared normalizer. The corrected estimator is a three-line change to standard implementations. An empirical consistency score built from the same measured relationships, not a re-derivation of the closed form, retrodicts 7 of 8 real dataset outcomes under standard (uncorrected) Group DRO, and on a real fairness benchmark the fix raises worst-group accuracy in dose-response with the minority's batch-absence rate ($r = 0.95$). The failure is invisible to any evaluation that only checks the metric the optimizer reports.
Chat is not available.
Successful Page Load