Adaptive Rank K-Subspace for Compressing LLMs
Samirasadat Jamalidinan ⋅ Yue Xu ⋅ Kazem Cheshmi
Abstract
Deploying foundation models on resource-constrained hardware requires reducing model size while preserving predictive quality. Existing low-rank compression methods typically rely on global approximations, complete rows or columns, predefined spatial blocks, or fixed basis-sharing patterns. We observe that Transformer weight matrices contain finer low-dimensional structure, where spatially distant tiles can often be represented by the same subspace. We introduce Adaptive-Rank K-Subspaces (ARKS), a post-training compression framework that treats fine-grained weight tiles as samples and uses subspace clustering to learn which tiles should share a low-rank basis, independent of their matrix locations. ARKS further allocates rank adaptively using a small calibration set and a global storage budget. Across OPT-125M and Llama-3.2-1B, ARKS reduces the storage of the targeted weight matrices by $53.3\%$ and $60.0\%$, respectively, with average relative perplexity increases of only $1.57\%$ and $10.13\%$ across three language-modeling benchmarks. At the same time, ARKS maintains competitive reasoning accuracy and outperforms the evaluated low-rank baselines at comparable storage levels.
Chat is not available.
Successful Page Load