MICA: Activation Checkpointing for Double-Backward Training of Machine Learning Interatomic Potentials
Hongyu Wang ⋅ Mingzhen Li ⋅ Hongtao Xu ⋅ Yuanchang Zhou ⋅ Weijian Liu ⋅ Weile Jia ⋅ Guangming Tan
Abstract
Conservative machine learning interatomic potentials (MLIPs) predict potential energy and obtain atomic forces as the negative gradient of energy with respect to atomic positions. Computing forces requires a backward pass from energy to atomic positions, and backpropagating the force loss differentiates through this pass, creating double-backward execution. This execution creates two activation classes with different lifetimes: forward activations and first-backward activations. Heterogeneous atomic graph batches make their memory footprint input dependent, while standard checkpointing mainly targets forward activations and does not directly manage first-backward activations. Therefore, we present MICA, an activation checkpointing planner for double-backward execution in conservative MLIP training. MICA combines a checkpoint primitive for double-backward phases with an input-aware dynamic programming planner that selects checkpoint modes for dynamic atomic graph batches under a specified memory budget. Across EqV3, MatRIS, PET, and UMA, MICA's full-checkpoint mode reduces peak memory by 49\% to 86\%, compared with only 10\% to 24\% from standard PyTorch checkpointing. On an 80GB GPU, training EquiformerV3 without checkpointing peaks at 75.47GB and exceeds 50GB in 76.7\% of 1000 iterations, while MICA keeps all 1000 iterations within a 50GB memory budget with 16.25\% step time overhead. Under the same GPU memory capacity, MICA increases the maximum feasible atoms per batch by $4.2\times$ to $9.2\times$. In distributed training, the reduced peak memory allows MICA to reduce reliance on heavier memory saving techniques, such as Fully Sharded Data Parallel (FSDP) and graph parallelism (GP). On a single H100 node with NVLink, MICA trains a MatRIS-MoE model with 2B parameters, achieving a $1.88\times$ to $2.03\times$ speedup over FSDP and GP baselines.
Chat is not available.
Successful Page Load