When Agent Responses Become Interaction State: Safety in a Deployed AI Teammate
Abstract
Safety evaluation typically asks whether an AI system responds safely to a harmful user input. But for agents that interact with users over many turns, a response does more than resolve the current request: it can shape the interaction that follows. We study this problem in a language-model-based, voice-enabled AI teammate deployed in a commercial multiplayer game. Using Korean and English interactions from the live service, we examine how safety failures emerge and persist across turns. We find that the main gap in the deployed safety specification was not a missing category of harmful content, but how the agent participated in the interaction. The agent could repeat unsafe speech, accept problematic relational roles, or make commitments that later resurfaced in otherwise ordinary gameplay. Based on these observations, we revise the safety specification to account for what an agent performs, accepts, and carries forward across turns. We then reconstruct production interaction histories and evaluate whether models can continue safely from them. When the agent has previously maintained the safety boundary, increasing user pressure has relatively little effect on harmless-response rates. In contrast, safety degrades sharply as earlier agent failures make up more of the inherited history, even for capable models given the revised specification at inference time. These findings show that long-horizon safety depends not only on responding safely to the current user, but also on the interaction state that an agent’s own behavior helps create.