AutoLMCompress: A Composable Framework for Autonomous Research on Language Model Compression
Abstract
We present AutoLMCompress, a verifiable, composable, and scalable framework for autonomous research on language-model compression. Human researchers define the research question, add their priors or inductive biases, and approve the evaluation protocol; within this scope, the framework reviews prior work, develops compression methods, evaluates them against baselines adapted to the target setting, and verifies the reported results. We apply AutoLMCompress to Per-Layer Embeddings (PLE) in Gemma 4, an instance of a recent trend toward scaling embedding capacity alongside Transformer computation. PLE accounts for 35.2% of the parameters in the E4B checkpoint, making it a substantial target for storage reduction. Motivated by cross-layer redundancy, the framework developed cross-layer latent vector quantization (VQ), which stores one quantized latent code tuple per token and maps it to each layer through a layer-specific affine decoder. Near 100× compression of PLE storage, cross-layer latent VQ achieves 97.8% retention across seven benchmarks (the geometric mean of per-benchmark performance relative to the same checkpoint with uncompressed PLE), compared with 66.3% for product quantization. With the official E4B checkpoint using 4-bit linear weights and 16-bit activations (W4A16), the same method achieves 96.9% retention. It outperforms the strongest evaluated baseline by 31.5 percentage points with BF16 weights and 14.0 percentage points with W4A16 weights.