Memorization Capacity and Grokking in LoRA Adapters
Abstract
LoRA adapters have become the canonical method for injecting new knowledge and behavior into a pretrained model, yet their physical limits are not yet fully understood. This work measures the capacity of adapters to both memorize and generalize by training on DataDecide OLMo checkpoints between 20M and 1B parameters. On random fact tables, we observe that an adapter stores about 1.2 bits per trainable parameter, a third of full training. Capacity follows a law conditioned on adapter size and base model quality and can be predicted on a held-out 1B base within 36\% error. Pretraining can raise the most capacity when base models are small, and capacity alone does not decide whether an adapter can generalize to rule-based facts. On modular addition, an adapter groks only above a fixed count of high rank and sufficient base quality. Finally, we propose a simple method for finding the minimum required rank a task required from one full-finetuning run, observing that real-world tasks require roughly 15 times less rank than modular rules.