Prefix Caching Can Silently Break Reproducibility in Interactive Agent Evaluation
Prathamesh Dinesh Joshi ⋅ Raj Dandekar ⋅ Rajat Dandekar ⋅ Sreedath Panat
Abstract
Interactive agents are scored by holding a conversation with a simulated user, so the resulting score can reflect variation from the agent, the simulator, the environment and the grader. We identify an additional source of variation at the serving layer. With both conversational models decoded greedily, 0 of 50 $\tau^2$-bench airline tasks and 0 of 50 retail tasks were fully reproducible across five replicates, and divergence usually began in the first generated message rather than accumulating over a long dialogue. In a matched $2\times2$ study over two engines and two cache settings, prefix-cache status predicted transcript repeatability exactly: 0 of 20 airline tasks reproduced with caching on, 20 of 20 with it off, on both vLLM and SGLang. Across matched cache-off arms, airline (20/20), retail (20/20) and telecom (11/11) reproduced exactly; banking improved from 4/19 to 7/19, with all 12 residual divergences first appearing at environment tool results, which localises what remains to the environment path without separating its language-model interface from retrieval tie-breaking. The benchmark's reward metric hides 76-84\% of these transcript divergences, which helps explain why the problem may go unnoticed. At two replicates per task, the estimated ordering of the variance components remains unstable. These results indicate that cache state should be reported alongside temperature and model version.
Chat is not available.
Successful Page Load