Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It
Jonathan Williams ⋅ Esin Tureci ⋅ Karthik Narasimhan
Abstract
Given a fixed dataset and reinforcement-learning-with-verifiable-rewards (RLVR) compute budget, should one train a single strong LoRA adapter, or split the same data and compute across $K$ independently trained shallow adapters? We compare the two allocations across four instruction-tuned models (1.5B-8B) and two domains (math and code) under exactly matched training data, training compute, and inference completions. On a five-benchmark math suite, adapter thickets - $K$ adapters trained on disjoint $\frac{1}{K}$ data shards with $\frac{1}{K}$ of the strong adapter's rollout budget each - win the majority vote in every (model, $K$) setting, by $+2.8$ p.p to $+10.3$ p.p at $K{=}16$. Weighting votes by response likelihood with a cross-validated temperature adds further points in every thicket setting while the single adapter gains nothing. Coverage is starker: at $160$-completions, the single adapter covers less than the untrained base model in seven of eight model - suite settings, while thickets exceed the base and the single adapter everywhere (by up to $+12.3$ p.p.). Mechanistically, thicket members are individually weaker but substantially less error-correlated than resamples of the single strong adapter, and both vanilla and likelihood-biased majority voting converts that diversity into accuracy. Under a fixed RLVR budget, breadth, not depth, is the better default.
Chat is not available.
Successful Page Load