QueueAgentBench: An Experimental Testbed for AI-Agent Queueing Formulation from Event Traces
Abstract
Operational queueing systems are often observed through event traces (e.g., arrival, service start, or departure) rather than through the explicit models required for downstream simulation and optimization. We study whether AI agents can formulate queueing models directly from such operational traces. To this end, we introduce \emph{QueueAgentBench}, a controlled testbed built on QGym that allows AI agents' queueing formulation performance to be studied systematically across variation in the amount and quality of operational traces, prompt guidance, and queueing structure. This testbed generates event traces from a representative set of queueing models, withholds the generating models from the agents, and evaluates the agents's formulations against ground truth. We demonstrate the application of QueueAgentBench through experiments with existing AI agents. The results show that AI-assisted queueing formulation is feasible but remains uneven: additional traces generally improve performance, corrupted traces can substantially reduce reliability, and prompt guidance is not uniformly beneficial. Performance also varies across queueing structures and agent systems, indicating that successful queueing formulation requires both identifying what must be inferred and accurately recovering it from the available traces. QueueAgentBench provides a controlled setting for studying the capabilities and limitations of AI agents in queueing modeling.