The Scaling Paradox of Tool-Calling LLM Agents Under Realistic MCP Faults
Antonio Ken Iannillo ⋅ Joshua S Owotogbe ⋅ Roberto Natella ⋅ Francesco Avallone ⋅ Kristina Kudryavtseva ⋅ Indika Kumara
Abstract
LLM agents increasingly rely on external tools exposed via the Model Context Protocol (MCP), yet current benchmarks assume tools respond correctly and on time, leaving agent resilience to realistic tool failures untested. We introduce MCP-Injector, a drop-in, protocol-aware fault-injection middleware that interposes on the MCP tool-call path with controlled persistence (permanent, transient, intermittent) and seedable schedules. Its fault model is empirically calibrated by mining 57,240 issue and pull-request threads across 1,774 open-source MCP repositories, yielding a distribution over six failure modes grounded in observed ecosystem incidents. In a 1,350-cell experiment crossing two backbone sizes (20B, 120B), three agent stacks (custom multi-round plan-and-execute, OpenAI Agents SDK, LangChain), and three fault injection conditions, we find that calibrated MCP faults cause severe quality degradation (60\% drop in task completion score, $p < 10^{-56}$) and reveal a scaling paradox: the larger model achieves higher fault-free scores but suffers lower completion rates under faults, driven by longer reasoning chains that amplify fault exposure. Framework-level error handling shapes resilience more than model scale, with stack choice producing significant interaction effects. These results demonstrate that capability benchmarks alone can mask operational brittleness, motivating protocol-aware fault injection as a complementary evaluation dimension for tool-calling agents.
Chat is not available.
Successful Page Load