Benchmarking LLM Agents for SOC Alert Triage: Performance and Failure Modes
Abstract
Tool-using large language model (LLM) agents are being explored for security operations, yet how different reasoning strategies perform and fail in end-to-end security operations center (SOC) alert triage remains unclear. A detector alert rarely contains enough context to resolve a case. The agent must retrieve surrounding telemetry and judge when it has enough information to dismiss the alert or escalate it to an analyst. An evaluation must therefore begin with the alert and examine both what the agent retrieves and when it ends the investigation. We introduce Alert-Bench, an interactive benchmark in which each investigation begins with one detector-generated alert and no predefined investigation question or preselected telemetry records. The agent searches surrounding telemetry through a fixed tool interface and returns a dismiss-or-escalate decision when it considers the investigation complete. The benchmark records the investigation trajectory. With the alerts, model, tools, and telemetry access held fixed, we compare five tool-using configurations in 6,235 investigations across 1,247 alerts from a four-day, multi-stage attack scenario. Every configuration misses at least 40.4% of attack-related alerts, and greater computational effort does not correspond to better triage. Attack-related alerts are dismissed more often in investigations with queries that return no records. For both agent configurations that include a revision stage, revision overturns more correct initial decisions than it repairs and leaves F1 lower. Yet only 26 of 225 attacks are missed by all five configurations, showing that failures differ across reasoning strategies. These findings indicate that alert-triage evaluation should examine investigation trajectories alongside final decisions.