MELD: Recursive Self-Improvement via Hierarchical Coordination for Multi-Agent Scientific Discovery
Abstract
Multi-agent scientific discovery with large language models raises a coupled learning problem: research assignments determine both which experiments are pursued and the experiences from which agents adapt. We introduce MELD, a unified hierarchical test-time training framework that jointly adapts a high-level Coordinator and multiple private Scientist policies within a single unseen discovery task. The Coordinator generates a joint portfolio of search branches and research directives over a shared derivation tree, while initially identical Scientists develop hypotheses and execution plans implemented by a shared frozen Executor. Both levels receive alternating GRPO-style parameter updates under a shared smooth best-of-budget objective. The resulting feedback loop enables recursive self-improvement at the policy level: coordination shapes Scientists' learning experiences, while their updated policies alter the outcomes that inform subsequent coordination. We establish objective alignment, characterize conditions for coordination gains over independent parallel discovery and the first-order benefit of private adaptation, and derive a finite-budget regret bound for adapted checkpoints. Across 12 MLGym tasks, MELD achieves the highest aggregate normalized scores among the evaluated systems, using a 4B backbone for coordination and research. The baselines use proprietary Scientist models, with all systems sharing the same frozen Executor.