Transferable, Time-Parallel Graph Dynamics via Space-Time Factorization
Wencong Yang ⋅ Chaopeng Shen ⋅ Abdolmehdi Behroozi
Abstract
Time-parallel neural operators (FNO and successors) have transformed PDE surrogates on regular grids. Mesh-based operators (MeshGraphNets, Transolver) extend operator learning to irregular geometries but remain autoregressive in time and trained one geometry at a time, while classical parallel-in-time methods converge poorly on nonlinear advection-dominated PDEs. We propose space-time factorization: partition the model into local temporal operators (learned, graph-independent) and a spatial mixer that propagates information across nodes without graph-dependent learned parameters. This structural condition restores time-parallelism on irregular graphs and yields cross-graph zero-shot transfer once three information leaks—graph-dependent weights, graph-specific features, and graph-correlated training distributions—are controlled. The same protocol applies to two distinct spatial-mixer families tested here—fixed-spectral and learned-attention—suggesting the principle is not tied to a single mixer choice. We validate the principle in three domains. On simulating floodwave propagation in river networks, a 27K-parameter model trained on the Yangtze transfers zero-shot to the Mississippi at $R^2=0.998$, training $28\times$ faster than a recurrent message-passing baseline and scaling to $5\times$ larger graphs. On cross-geometry Navier–Stokes, both Chebyshev (SepFNO) and Physics-Attention (SepTransolver) variants reach within ~10% (relative $L_2$) of a locally-trained oracle using only 15 training geometries—to our knowledge the first demonstration of knowledge accumulation in this regime across fundamentally different topologies. On cross-city traffic forecasting, zero-shot transfer reaches within 12.9% of an oracle trained on the target city. A consistent pattern emerges across these settings: cross-geometry error decreases with both the number of training geometries and their relatedness to the test, with the rate of accumulation depending on the spatial mixer. The result opens the path for pooled-data, foundation-style models for graph dynamics on irregular topologies.
Chat is not available.
Successful Page Load