Who Pays for AI, and Whose Language Does It Speak?
Abstract
Data centres are increasingly built in Southeast Asia. The benchmarks used to certify the models they train are built elsewhere. We measure both on a single axis. Using Ember grid emissions data for 2025 and per-site power usage effectiveness from Google's 2026 Environmental Report, we compute the carbon cost of a kilowatt-hour of IT load across nine data centre regions. Southeast Asian sites carry 2.0 to 2.8 times the carbon cost of Dublin, for instance. Grid composition drives this difference, while the cooling penalty usually assumed for tropical climates accounts for 4 to 7 percent of it. We then audit FLORES-200, Belebele, and XTREME against the languages of those regions. Every language we examined appears in all three benchmarks, so coverage alone does not explain the gap. Each language is covered in a single script, all items are translation-derived, and no contact variety appears in any benchmark. Malay and Javanese, spoken where compute is most carbon-intensive, satisfy none of three validity criteria we define. We argue that the problem is one of measurement validity rather than inclusion, and we propose that models report the language profile of the regions supplying their compute.