Behavior Beyond Text: Evaluating Refusal Transfer Failure in LLM Agents
Abstract
While large language models are more frequently being deployed in the form of tool-calling agents in high-stakes settings, their safety evaluations remain heavily concentrated on text refusal behavior. To what degree refusal mechanisms learned during text alignment transfer to tool calls remains largely untested. We begin by constructing an evaluation dataset of 2304 prompts that cover four domains and are grouped into triplets corresponding to a single harmful request, including a no-tool prompt, a standard tool-enabled prompt, and an adversarially framed tool-enabled prompt. We objectively measure tool call safety using 20 pre-defined forbidden actions. Upon evaluating five language models of varying sizes and model families, we observe a divergence between refusals in text and refusals in tool calls. Following the behavioral analysis, we then mechanistically investigate the divergence. First, we identify a linear refusal direction in the residual stream and find that this signal reliably predicts unsafe calls across all five model families. However, activation patching only partially restores refusals, and what it restores is a text refusal in place of the call rather than a safer call, showing the direction is a partial mediator of whether the model acts, not of how safely it acts; the remaining gap is not identified by our interventions and may involve mechanisms beyond this single direction. Furthermore, when steering the model using the refusal direction, we find that it suppresses tool use on harmful and benign prompts alike rather than selectively removing unsafe calls. The refusal direction suppression we observe in tool calls across most of the tested architectures suggests that LLM agents in high-stakes settings will require separate evaluations for tool safety and text safety.