User-in-the-Loop World Simulation for Evaluating Streaming Interactive Agents
Abstract
Interactive assistants differ fundamentally from passive perception models in that their outputs can influence subsequent user actions. A well-timed intervention may prevent an impending error, whereas misguided advice may induce a failure that would not otherwise occur. The effectiveness of an assistant response can therefore be determined only from the user behavior and task outcome that follow from it. Evaluating an interactive assistant consequently requires executing the user actions elicited by its responses and measuring their effects on the task outcome. We introduce a simulation-based, response-contingent evaluation framework for streaming interactive assistants. Because a language assistant influences a task through the user’s decisions, the framework includes a responsive simulated user that interprets the assistant’s guidance and determines how to act. The task environment executes the resulting action and renders its consequences as the next observation, allowing different assistant responses to produce different rollouts and outcomes. We implement the framework in three simulation engines spanning cooking, household manipulation, and mobile interface use. The engines support persona-conditioned users and controlled instances of errors, questions, and plan deviations, while their execution traces enable objective outcome-based evaluation metrics. As a first instantiation of the framework, we programmatically construct a 231-case benchmark and evaluate state-of-the-art multimodal models across more than 2,000 rollouts. The strongest model completes only 36.1\% of tasks within the time limit and exhibits poor intervention timing and low communication efficiency, revealing a substantial gap toward reliable interactive assistance.