WaterPrune: Waterfilling Sparse Foundation Models in Resource-Constrained AI
Abstract
Access to foundation models is shaped by data, language coverage, and the cost of storing and serving the models. Post-training pruning offers a training-free route to smaller deployments, yet uniform sparsity rules spend the pruning budget unevenly across layers. We study layer-wise sparsity allocation as a resource-aware optimization problem. For each layer or module, we measure a direct terminal-loss curve under SparseGPT-style one-shot pruning, convexify the measured curve, and solve a waterfilling allocation problem under a global sparsity budget. The resulting system, WaterPrune, is label-free, calibration-only, and task-training-free. Across high-sparsity language-model and zero-shot evaluations, WaterPrune improves over uniform pruning and evaluated layer-allocation competitors in the regimes where compression pressure is strongest. These results support high-sparsity pruning as a practical path toward resource-constrained foundation model deployment.