When Must AI Training Stage Checkpoints? A Distributional Model of Durability Boundaries at Scale
Farid Talibli ⋅ Awais Khan ⋅ Christopher Zimmer ⋅ Sungyong Park ⋅ Jihoon Yang ⋅ Youngjae Kim
Abstract
Large-scale AI training runs span days within fixed HPC allocations, making checkpointing essential for bounded recovery after failure. Asynchronous checkpointing hides checkpoint latency from the training loop, but does not guarantee that the most recent state has reached a durable target: data may still be in flight to the shared parallel filesystem (PFS) when a failure occurs, forcing rollback to an older checkpoint. We formalize this durability gap through a distributional model that combines hardware crossover analysis with statistical models of PFS completion time under scale and contention. Rank-local completion times follow a lognormal body distribution, while wall-clock completion is governed by an extreme-value process determined by the slowest writer. The resulting analysis yields a contention-dependent threshold,~$C^\star_p(n)$, beyond which the PFS is no longer the fastest path to durable persistence. Guided by this analysis, we design CacheX, a transparent staging layer that composes with PyTorch Distributed Checkpoint~(DCP). CacheX stages bulk data to local DRAM and replicates it to a neighboring node's NVMe, providing allocation-local durability while preserving eventual persistence to the PFS. On Frontier, across IOR runs up to 256~nodes and end-to-end FSDP+DCP training up to 64~nodes, CacheX reduces write-time variability by up to~$43\times$, achieves a replica-durable effective throughput of~${\sim}\,2.5$ to $2.8$\,GiB/s per node decoupled from PFS contention, and reduces worst-case rollback exposure. Our results show that staging is not a universal accelerator; rather, it is a conditional policy for selecting the first durable checkpoint boundary under measured scale, load, and tail behavior.
Chat is not available.
Successful Page Load