Toward a Multimodal Foundation Model for Plasma State Reconstruction
Bin Xia ⋅ Chuanfei Dong ⋅ Boting Li ⋅ Yilan Qin ⋅ Zikang Xie ⋅ Sung H Son ⋅ Adam Stanier ⋅ Jongsoo Yoo ⋅ Hantao Ji ⋅ Ahmed Diallo ⋅ Han Ding ⋅ John Wise
Abstract
Plasma experiments observe coupled physical states through diagnostics that are sparse, heterogeneous, and incomplete in both space and time. We present a proof-of-concept toward a multimodal foundation model for plasma state reconstruction during magnetic reconnection, motivated by the diagnostic environment of the Facility for Laboratory Reconnection Experiments (FLARE). The model is trained on 250 VPIC magnetic reconnection simulations at fixed plasma beta, with 225 complete runs for training and 25 held-out runs for evaluation. A single 5.5M-parameter residual 3D U-Net receives masked $B_x,B_y,B_z$, and density fields together with explicit observation masks, preserves temporal resolution through spatial-only pooling, and uses global spatiotemporal self-attention with rotary positional encoding. Training independently masks the magnetic vector and density modalities with spatially random, grid, block, and temporal missingness over a wide range of observation fractions. The resulting model supports multiple reconstruction settings without task-specific retraining. When density is completely hidden, as little as 0.1% magnetic spatial visibility reduces median density NRMSE from approximately 0.9 to 0.1; conversely, fully observed density substantially improves current-density reconstruction when magnetic observations are absent or extremely sparse. With no magnetic observations, roughly 1% density visibility is sufficient to approach the density-reconstruction error obtained with fully observed magnetic fields, although accurate current structure remains much more dependent on direct magnetic measurements. Full magnetic conditioning also stabilizes conditional completion of missing future density frames. These results demonstrate bidirectional cross-modal information transfer and support masked multimodal reconstruction as a promising route toward reusable plasma state-estimation models, including future FLARE applications. The present study is limited to synthetic VPIC data at one plasma beta and does not yet establish sim-to-real or cross-regime generalization.
Chat is not available.
Successful Page Load