Load-Shedding Hurts Some Languages More: Outage Structure and Language Equity in Distributed Training
Abstract
Multilingual models are increasingly trained on infrastructure inside the regions whose languages they serve, where scheduled load-shedding takes whole regions offline for hours at a time. We ask how the structure of those outages shapes which languages the model ends up serving well. An outage rotation can be participation-fair — every region offline for exactly the same total time — leaving each region's share of the training data unchanged while still producing unequal quality across the languages those regions host. Holding total downtime fixed and varying only outage length, the cross-language loss spread grows from 0.09 to 0.45 as outages lengthen from one round to sixteen, while a control that shuffles each worker's on/off sequence in time (preserving uptime exactly) shows no such trend. The effect persists on multilingual mC4, cannot be attributed to cumulative data imbalance (zero by construction), and largely disappears when the server momentum is removed. Repairs that reuse a region's own last update help less the longer the outage, and we show why: an update direction has a measurable shelf life (cosine 0.80 to 0.08 over 16 rounds) while a similar online region's current update holds at 0.70 to 0.66. Substituting present neighbours instead cuts the gap from 0.61 to 0.14 at the longest outages, and weighting that substitution by the online-estimated shelf life removes the last tuned constant. Equal access to the grid is therefore not equal access to the model: the temporal structure of availability, not only its total, shapes which languages a shared model learns well.