T-ReX: Learning Tile-Reuse Indexes for Structured Model Compression
Abstract
The rapid proliferation of large-scale machine learning models, such as LLMs, has driven remarkable progress across application domains. Nevertheless, their scale comes with significant deployment challenges and inherent redundancies that remain open challenges to this day. To this end, we propose the Tile-Reuse Index (T-ReX), a structured model compression method that identifies and indexes reusable, contiguous parameter groups (tiles) across a model's architecture. By learning a compact set of matching tiles, we compress the model's memory footprint while preserving its performance. Furthermore, by reconstructing dense layers on the fly at inference time, T-ReX compressed models remain efficient at runtime. In our experiments, we show that T-ReX outperforms existing compression methods across benchmarks by most closely matching the original model's performance. We furthermore show how post-training on top of T-ReX enables small footprint models to recover most of the performance loss introduced by compression.