Modular Norm RandOpt: Population-Efficient Ensembling through Architecture-Aware Perturbations
Kirato Yoshihara ⋅ Hiroaki Hamade
Abstract
RandOpt adapts a pretrained language model by sampling $N$ weight-perturbed candidate models, ranking them on a small selection set, retaining the $K$ highest-reward candidates, and combining their held-out predictions by plurality vote. Its isotropic search applies a single coordinate scale to parameter tensors with different shapes and functions. We introduce \emph{Modular Norm RandOpt} (MN-RandOpt), which modifies only candidate generation: each tensor perturbation is normalized by a role-specific norm and rescaled using a fixed profile derived from model architecture and source-task sensitivity calibration. On Qwen2.5-1.5B-Instruct, MN-RandOpt exceeds RandOpt using $3\times$ fewer candidates on Countdown and at least $12\times$ fewer on GSM8K. Across seven tasks and three Qwen model sizes, MN-RandOpt improves over RandOpt in $14$ of $21$ settings and ties in one additional setting. These results show that architecture-aware perturbation geometry can substantially improve RandOpt's population efficiency while preserving its ranking and ensembling procedure.
Chat is not available.
Successful Page Load