Hops Are Not Distances: The Halo Cost of Partitioning Multi-Resolution Weather Graph Neural Networks
Abstract
Large flat mesh graph neural networks are split across processors, with each processor holding a partition of the mesh and a halo of neighbouring nodes, so that the partitioned computation is arithmetically identical to the un-partitioned one. With L message-passing layers the halo must be at least L deep to assure exactness. In multi-resolution meshes, due to their 3-dimensional nature, a message-passing layer expands the dependency set across resolution levels as well as within one, which grows the halo required for exactness far faster. We compute that dependency set by reversing the message-passing schedule and measure it across the architecture families used for weather forecasting. A GraphCast-style multi-mesh saturates after 7-9 message-passing layers, although the model runs 16; at that depth an exact static halo approaches whole-mesh replication on every rank. Exact spatial partitioning of multi-resolution GNNs is therefore constrained by inter-resolution dependency growth, which limits the scalability benefits of partition-based parallelism.