Chat-safe is not CUA-safe: Chat methods may not transfer to computer use
Abstract
Safety of multi-modal large language models in the computer-use agent (CUA) setting depends on the structured function calls emitted by the model that drives mouse and keyboard actions, not purely on the model's natural language outputs. However, most safety evaluations still measure harm refusal in chat-only settings, where the model only produces a text reply. We ask whether chat safety transfers to the agentic CUA setting, and whether white-box interventions that improve chat refusal remain effective when the model must generate computer tool calls. We evaluate four open-weight models and four closed frontier models on harmful desktop tasks from OS-Harm in text-only chat versus multi-turn agent mode in a desktop environment. Harm refusal drops sharply from chat to agent mode on Holo3 and Claude; GPT refusal changes little. We show that attempted harmful actions occur even when task completion is rare, implying that low task success is not evidence of safety. Applying activation steering with a harm direction to the open-weight models, we find that its improvement of chat refusal has little selective effect on agent-mode behavior and can reduce success on harmless computer-use tasks. Our results serve as a warning: safety claims for CUAs have to be checked in the agentic setting on verbal refusal, attempted harmful action, and successful task completion.