Benchmarking Small Language Models as Agentic Workflow Executors in Enterprise Customer Support
Abstract
Enterprise agents must do more than produce plausible text: they must execute bounded, stateful workflows while respecting tool, policy, and conversational constraints. In this paper, we introduce a deployment-grounded benchmark for evaluating skill executors, specialized language models operating inside a selected enterprise workflow. Starting from approximately 350K real customer-support transcripts, our pipeline mines 858 workflow clusters, validates 130 structured skills, and constructs a fixed evaluation set of 500 multi-turn scenarios supported by 281 Model Context Protocol (MCP)-compatible mocked tools. We evaluate 17 closed and sub-20B open-weight small language models (SLMs) zero-shot with a deterministic checker and five independent LLM judges. GPT-5.4 obtains the highest mean judged pass rate (45.0%), while QWEN3.5-9B is the strongest sub-20B open model (21.6%) and exceeds several closed baselines. Judge choice materially changes absolute scores: mean pairwise agreement is 84.2%, but Cohen’s κ is only 0.524.