WURI: Watching Unfolding Risk in Agent Interactions
Abstract
Large language model (LLM) agents increasingly operate through multi-step tool-use trajectories, where harmful intent may be distributed across actions that appear benign in isolation. This challenges single-step moderation, while repeated LLM-as-judge evaluation over growing prefixes can be costly and may intervene too late. We introduce WURI (Watching Unfolding Risk in Agent Interactions), a lightweight monitor for early prefix-level detection of harmful agent trajectories. WURI encodes each observed textual step with a frozen text encoder, learns an atom-adapted trajectory representation through metric learning, and scores each prefix with a fixed prototype-margin rule. It requires no access to agent model weights or hidden states, does not modify the agent's internal model, and uses no external LLM-as-judge calls at inference. Across five generalization settings, WURI achieves the best average prefix-area under the detection curve (AUDC), ranks first on three settings, and reaches the strict early-detection operating point within four steps in multiple evaluation settings. Runtime analysis shows a wall-clock speedup of over two orders of magnitude compared with repeated full-prefix guardrail evaluation. These results demonstrate that representation-based prefix scoring is an effective and efficient direction for monitoring unfolding risk in LLM agent interactions.