KV-Pipe: On the Relation Between KV Sharing and Pipeline Parallel Efficiency in LLMs
Hossein Rajabzadeh ⋅ Maryam Dialameh ⋅ Harish Krishnamoorthy Murali ⋅ Walid Ahmed ⋅ Weiwei Zhang ⋅ HYOCK JU KWON
Abstract
Pipeline parallelism (PP) is widely used to scale large language model (LLM) training, but its efficiency is often limited by stage imbalance and pipeline bubbles. Cross-layer KV sharing has primarily been studied for reducing KV-cache costs during inference, without examining how KV reuse reshapes pipeline workloads. We present \textbf{KV-Pipe}, a stage-aware KV-sharing mechanism that turns KV reuse into a pipeline-balancing control knob. KV-Pipe starts from the tail stage, converts selected attention layers to cross-layer KV sharing, and iteratively retargets the current bottleneck to drive the FLOPs Imbalance Ratio (FIR) toward $1$. The procedure runs offline using only the pipeline partition and per-layer FLOP estimates, requiring no online tuning. Across multiple pipeline-parallel configurations, KV-Pipe achieves up to \textbf{9.2\%} relative MFU improvement and a \textbf{9.8\%} reduction in iteration time over the 1F1B baseline, with larger gains at higher PP degrees. An inference case study further shows that the same KV-sharing mechanism reduces KV-cache growth and redundant KV projection work, improving long-context decoding throughput. These results identify KV sharing as a system--architecture degree of freedom for jointly improving pipeline-parallel training efficiency and long-context inference.
Chat is not available.
Successful Page Load