When Lightweight Routers Fail: Distribution Shift and Deferral in Tool-Calling LLM Agents
Abstract
Lightweight LLM routers can reduce the cost of tool-calling agents, but their reliability can deteriorate when deployment differs from the training distribution. We evaluate embedding-based lightweight routers on the Berkeley Function-Calling Leaderboard (BFCL). We show that (1) simple validity checks on generated tool calls provide an inexpensive runtime verification signal, recovering 63\% of Haiku's performance relative to Nano alone with only a 12\% cost increase. (2) Routers that perform well in-distribution degrade under held-out task-family shift; we observe similar limitations when transferring single-step-trained routers to multi-step settings. (3) We propose a cascading router to defer queries for which no available models has an observed success; at best, the cascading router achieves 88\% precision but only 13\% recall for deference detection. Together, these results indicate that lightweight routers can reduce cost on familiar in-distribution tasks, but they should not be treated as reliable failure detectors especially when the task distribution changes. In many cases, deterministic validity checks naturally available in tool-calling settings provide a cheap and reliable complementary signal for detecting failures and triggering escalation.