MARGENT: Measuring the Marginal Value of Delegation in Agentic Systems
Abstract
Current agentic systems increasingly rely on multiple specialized agents to improve performance, but each additional interaction requires more model computation and introduces new failure paths that may leave the solution unchanged or even corrupt a correct answer. Many existing systems follow fixed interaction patterns or, when selective, route from query-level features or model confidence; these signals do not directly compare the outcomes of committing and delegating from the same reasoning state. This raises a fundamental question for agentic systems: When is additional intelligence worth it? We study this question in a concrete multi-agent setting, where a manager must either commit to its current candidate solution or delegate to another sub-agent. We introduce MARGENT, a trainable framework that branches from the same manager-generated reasoning state into immediate commitment and delegation to each available sub-agent. By comparing their final-answer correctness, MARGENT measures the marginal value of each delegation and distills the shortest successful branches into supervised trajectories for policy learning. Across four reasoning benchmarks, the best single delegation recovers 91.7% of the total accuracy gain available under the measured oracle on average. On four collection-disjoint evaluation sets, the final system exceeds the same trained manager's pre-delegation accuracy by 17.1 percentage points while making 0.53 sub-agent calls per example. On development diagnostics, with roughly one selected call instead of three, it also matches forced-all delegation on MMLU-Pro (71.5%) and exceeds it on AQuA-RAT (81.1% vs. 75.2%). These results show that effective agentic systems depend not on invoking more agents, but on allocating additional intelligence according to its marginal value at the current reasoning state.