On the Geometry and Latent-Space Composition of Hypernetwork-Generated LoRAs
Abstract
Recent work proposes hypernetworks that convert a document into a low-rank parameter update (LoRA) in a single forward pass, enabling fast injection of new knowledge into a frozen LLM. However, storing separate LoRA adapters for each document can be unsustainable in streaming settings. In this paper, we study the problem of composing hypernetwork-generated LoRAs using two representative systems, Doc-to-LoRA and SHINE. We show that hypernetwork-generated LoRAs have a geometry sharply different from directly fine-tuned LoRAs. Hypernetwork-generated LoRAs are highly spectrally concentrated and strongly aligned, while next-token prediction and QA fine-tuned adapters are diffuse and nearly orthogonal across documents. This suggests that hypernetwork-generated LoRAs are not arbitrary low-rank weight updates, but decoded points on a structured adapter manifold. Guided by this observation, we propose latent-space LoRA merging: instead of merging in the weight space, we compose the intermediate hypernetwork latents and then decode the merged latent through the frozen LoRA-generation head. We show that latent-space merging techniques consistently improve over weight-space and factor-space baselines while preserving a single fixed-rank adapter. These results establish latent-space composition as a promising mechanism for bounded accumulation of hypernetwork-generated LoRAs in long document sequences.