My Budget Is Yours: Learning When to Hand Off in LLM Planner-Executor Teams
Abstract
Multi-agent systems of large language models (LLMs) solve complex tasks by dividing the work across cooperating agents. A common instance is the planner-executor pattern, where one agent reasons out a plan and another implements it. When such a team works under a single token budget, every token the planner spends deliberating is a token the executor loses. However, in existing systems this allocation is imposed from outside, by external orchestrators or fixed schedules, never by the deliberating agent itself. We present L2H (Learning to Hand Off), a planner trained to make this endogenous stopping decision by learning to time its own handoff, balancing its deliberation against its teammate's execution needs. We train L2H in two stages: cold-start supervised fine-tuning (SFT) on budget-conditioned teacher traces establishes the deliberate-then-commit behavior, and peer-aware reinforcement learning then optimizes the handoff against the team's joint outcome. Through this signal, the planner learns cooperative deliberation: predicting its peer's needs and stopping while enough budget remains. Across five code benchmarks and five total token budgets, L2H attains the best average pass rate, exceeding even the best fixed split chosen in hindsight. These results suggest that cooperative budget governance from within the team can replace external orchestration, once the deliberating agent learns to spend with its teammate in mind.