PipeFSDP: Efficient Pipeline Parallel under Fully Sharded Data Parallel for Large Language Model Training
Xinglin Pan ⋅ Mingji Han ⋅ Penghao Zhao ⋅ Lin Zheng ⋅ Rayying ⋅ key ⋅ Shaohuai Shi ⋅ Xiaowen Chu
Abstract
It is now a common practice to use PyTorch’s official Fully Sharded Data Parallel (FSDP) framework to train large language models (LLMs) on large-scale GPU clusters, typically together with pipeline parallelism (PP) to achieve better performance. However, existing popular systems like TorchTitan with FSDP and PP are inefficient due to the large bubbles and peak memory footprint caused by suboptimal task scheduling. In this paper, we propose PipeFSDP, an efficient training framework for FSDP combined with PP that alleviates pipeline bubbles and resource contention, thus improve training efficiency. Specifically, we design four fine-grained optimizations: dynamic context-aware prefetching, communication contention mitigation, redundant reshard elimination, and communication order optimization. In addition, we develop a heuristic model that selects the FSDP and PP degrees based on the given hardware and model configurations, thereby enhancing the practicality of PipeFSDP. Empirical evaluations on both dense and sparse LLMs (Llama3, Qwen3, and Qwen3-MoE) across 32-GPU and 64-GPU clusters demonstrate that PipeFSDP significantly outperforms the Torchtitan system by up to 1.28$\times$ in Model FLOPs Utilization (MFU) and 1.87$\times$ in memory consumption.
Chat is not available.
Successful Page Load