Config2Weights: Architecture-Conditioned Generative Initialization of LLMs
Abstract
Training large language models (LLMs) from scratch is prohibitively expensive, motivating methods that reuse parameters from pretrained model weights such as weight generators which are largely limited to non-LLMs full weights or low-rank adapters for LLMs, and model expansion which depends on architecture-specific transformations from a chosen parent. Neither yields an initialization for an arbitrary target architecture. We introduce Config2Weights (\ours), an architecture-conditioned framework that synthesizes LLM initialization weights from pretrained model zoos. \ours represents heterogeneous parameters as fixed-size weight chunks paired with structural and architectural metadata, so that a single generator covers multiple decoder-only architectures. Trained on this representation, \ours learns the distribution of pretrained weights given architecture and samples an initialization for a new model from its configuration alone. Experiments show faster convergence than random initialization and competitive performance with specialized width- and depth-expansion methods, without explicit parent-to-target transformation rules.