From Scaling Tests to Production: Dense LLM Pretraining on Leonardo
Abstract
Efficient large-scale LLM pretraining requires jointly selecting global batch size, GPU count, and parallelism configuration, since configurations that maximize hardware efficiency may not be compatible with convergence requirements. We study this trade-off on the Leonardo supercomputer using OpenEuroLLM Prelude, a 9B-parameter dense model. Under weak scaling, per-GPU computational efficiency remains approximately constant at 50\% up to 1024 NVIDIA A100 GPUs, indicating limited scale-dependent degradation when the workload per GPU is preserved. However, under strong scaling with a fixed global batch size of 1024, efficiency declines beyond 128 GPUs and reaches 30\% at 1024 GPUs, showing that increasing GPU count becomes ineffective once the workload per GPU is too small relative to distributed-training overheads. Increasing the global batch size recovers a substantial fraction of this loss, with GBS=4096 providing the highest hardware efficiency, but GBS=2048 is selected for production because it remains within the range supported by the available convergence-oriented scaling-law estimates. At this batch size, optimizing the parallelism configuration increases efficiency from 36.9\% to 44.2\% at 1024 GPUs. The resulting production run achieves a mean efficiency of approximately 42.5\%, despite substantially greater step-to-step variability than the controlled benchmarks. These results show that short scaling experiments, when combined with convergence constraints, provide a practical basis for jointly selecting batch size, GPU allocation, and parallelism configuration, while also giving a useful estimate of sustained production throughput.