ROCKET: Rapid Optimization via Calibration-guided Knapsack-Enhanced Truncation for Efficient Model Compression
Ammar Ali ⋅ Baher Mohammad ⋅ Denis Makhov ⋅ Dmitriy Shopkhoev ⋅ Magauiya Zhussip ⋅ Stamatios Lefkimmiatis
Abstract
We present $\textbf{ROCKET}$, a training-free model compression method that achieves state-of-the-art performance in comparison with factorization, structured-sparsification and dynamic compression baselines. Operating under a global compression budget, ROCKET comprises two key innovations: First, it formulates layer-wise compression allocation as a multi-choice knapsack problem, selecting the optimal compression level for each layer to minimize total reconstruction error while adhering to a target model size. Second, it introduces a single-step sparse matrix factorization inspired by dictionary learning: using only a small calibration set, it sparsifies weight coefficients based on activation-weights sensitivity and then updates the dictionary in closed form via least squares bypassing iterative optimization, sparse coding, or backpropagation entirely. $\textbf{ROCKET}$ consistently outperforms existing compression approaches across different model architectures at 20–50\% compression rates. Notably, it retains over 90\% of the original model’s performance at 30\% compression without any fine-tuning. We will release the code implementing ROCKET upon paper acceptance.
Chat is not available.
Successful Page Load