Continual Learning at Fixed Update Bandwidth: Budgeting Quantization-Cell Exits in Deployed Models
Shrenik Bhansali ⋅ Larry Heck
Abstract
A deployed language model is a low-bit quantized integer model, replicated many times over, so every continual update must be serialized and shipped to every replica. This is a recurring cost that continual-learning research does not measure, and it scales with how often a model adapts rather than with how much it learns. We observe that under a \emph{frozen} deployment grid a group-wise quantizer partitions weight space into cells, and that moving a weight inside its cell leaves the shipped integer model bit-identical. The update that must actually be transmitted is therefore exactly the set of changed codes, which makes update bandwidth something a regularizer can act on directly. We propose Cell-TR, which budgets cell exits during continual fine-tuning and is evaluated by the size of the serialized code patch rather than by a proxy. Conventional regularizers do not reach this regime at any strength we tested: sweeping the EWC coefficient across three orders of magnitude saturates well short of it, and replay's payload is nearly invariant to its ratio. Cell-TR instead reaches 3.71~MB at accuracy .879, $14.0\times$ smaller than sequential fine-tuning with no detected loss of accuracy. Two ablations locate the effect in what the penalty is anchored to rather than in the shape of the barrier: widening the free region inside a cell only costs bytes, and against the current-bin-center rule of quantization-aware training, anchoring instead to the previously \emph{deployed} state is cheaper in 20 of 20 matched configurations. The advantage grows with update frequency, which is what makes it relevant to models expected to keep learning after deployment.
Chat is not available.
Successful Page Load