Learning to Route Safely in Multi-Agent LLM Systems
Abstract
Multi-agent LLM systems (MAS) coordinate multiple model calls, deciding which models see a query, which agents are instantiated, and how their outputs are aggregated. Learned routers cast these decisions as a policy trained to maximise clean accuracy, but routing for accuracy alone leaves three failure modes: harmful responses to jailbreak prompts, wrong final answers when some routed agents are replaced by malicious agents, and leakage of confidential context. Existing defences either moderate the response after generation, leaving the underlying route in place, or search for a safe scaffold offline without training the route itself under multiple safety signals. We formulate MAS routing as constrained reinforcement learning, optimising jailbreak safety and byzantine utility under malicious-agent replacement with an explicit clean-accuracy floor enforced by an adaptive Lagrange multiplier. Across controlled evaluations, a single shared router preserves clean accuracy while improving all evaluated safety metrics, Pareto-dominating the evaluated multi-agent baselines; the ordering persists under an independent safety judge and malicious-agent replacement rates from 25\% to 75\%.