Transmuting Prompts Into Weights
Abstract
Large language models (LLMs) achieving breakthrough performance in long-context regimes face severe computational and memory bottlenecks when processing extensive prompts, multi-turn histories, and long-horizon instructions. While inference-time control techniques exist, they often rely on empirical heuristics. Recently, Dherin et al. demonstrated that the conditioning effect of a prompt can be mapped to token-dependent implicit weight updates, introducing static thought patches for prompt compression. To address the scaling challenges of long-context foundation models, we extend this framework into a formalized model editing algorithm and derive a principled method for efficiently condensing extensive prompt contexts into lightweight, token-independent thought vectors and matrices. By transmuting redundant long-context computations into compact parameter-space updates, our approach bypasses heavy sequence processing at inference time, offering a resource-efficient efficiency technique for long-context reasoning, context compression, and dynamic knowledge injection in demanding downstream applications.