Planning Under Lock-In: Multi-Objective Hierarchical Reinforcement Learning for Long-Horizon AI Data Center Build-Out
Satyapragnya Kar ⋅ Rajarshi Das
Abstract
The AI data center build-out is among the largest infrastructure undertakings in history: global data center capital expenditure is projected to reach roughly \$7\,trillion by 2030, with more than \$700\,billion already committed through 2027, annual electricity demand heading toward 945\,TWh, or close to 3\% of world power, site water demand roughly tripling, and sector carbon footprint roughly doubling over the same decade [1--8]. The binding constraint is no longer silicon but energized capacity: approximately 2{,}060\,GW sits in U.S. interconnection queues facing three to seven year waits, against 40\,GW of new data center demand announced in 2024 alone [3, 6]. At an all-in cost reported in the tens of millions of dollars per megawatt for AI-capable capacity, a single build phase is a multi-billion-dollar commitment [4]. Yet the decisions that shape this build-out (where to site a facility, how to procure power, which cooling architecture to install, and when to expand) are overwhelmingly made with static single-point scoring tools. Such tools are poorly matched to the problem. A site's water, power, and carbon profile drifts substantially across a ten to twenty year asset life, and the decisions themselves are irreversible, binding capital and physical plant for years after they are taken. We formalise long-horizon build-out as a tiered multi-objective Markov decision process in which strategic commitments are gated to discrete windows, carry switching costs, and energise only after source-dependent lead times, while operational dispatch is restricted to capacity that physically exists. We solve it with a preference-conditioned multi-objective agent augmented by a semi-Markov manager and worker hierarchy, recovering a family of policies that a planner can steer across the cost, carbon, water, and reliability trade-off surface at inference time. Across a three-arm irreversibility ablation on six sites and four demand and climate scenarios, we find that preference conditioning produces steerable trade-offs with correlation magnitudes up to $0.92$ on the water and reliability axis; that imposing lock-in improves both frontier spread and held-out generalisation relative to a freely reversible formulation; and that learned policies adapt several years ahead of a binding water-stress threshold. We release the environment, agents, baselines, and evaluation harness to support further work on anticipatory infrastructure planning.
Chat is not available.
Successful Page Load