LoRA on a Budget: Function-Aware Post-Hoc Compression for Adapter Fleets
Abstract
Low-Rank Adaptation (LoRA) has become a key mechanism for efficiently specializing foundation models into many task- and user-specific variants. Serving many LoRA adapters under a shared resource budget turns post-hoc compression into a two-level allocation problem: rank must be distributed across layers within each adapter and across adapters within the fleet. Although trained LoRAs contain substantial spectral redundancy, spectral reconstruction alone does not determine how much compression an adapter can functionally tolerate. We formulate fleet compression through a functional rank-distortion frontier for each adapter, followed by fleet-level minimax allocation under a shared rank budget. Given exact frontiers, this decomposition is equivalent to joint optimization over all layer ranks. We introduce Functional Rank Allocation (FRA), an unlabeled method for constructing practical frontiers. FRA measures isolated module truncations end to end, uses these measurements as an additive search surrogate solved by dynamic programming, remeasures every complete candidate, and takes the lower-risk envelope with a complementary spectral proposal before exact fleet allocation. Across 78 adapters from three public fleets, FRA consistently improves worst-case and tail utility under constrained rank budgets, with the largest gains appearing when functional compression sensitivity is heterogeneous across adapters. Its advantage depends on how well isolated truncation effects compose, identifying compositionality as a measurable failure mode of functional inner allocation.