Evaluating the Robustness of LLM Cybersecurity Agents to Adversarial Telemetry in Security Logs
Abstract
Security teams are increasingly deploying AI agents to assist human analysts with alert triage and incident investigation. This creates a vulnerability when attackers can place natural-language text in the security telemetry these systems reason over, potentially causing an agent to misclassify malicious activity, prematurely close an alert, or divert an investigation, thereby allowing an intrusion to persist undetected. We study whether indirect prompt injection (IPI) embedded in security telemetry can cause an LLM security agent to downgrade malicious activity, and whether access to additional evidence reduces this risk. We construct security evidence from realistic multi-stage attack chains and compare model judgments on matched clean and poisoned inputs. Across five model configurations and more than 1,100 prompt injection payloads spanning four payload framings, telemetry-based IPI causes verdict downgrades in every evaluated configuration. Larger models and thinking-enabled variants are not consistently more robust, and attacks that mimic trusted authority are particularly effective. We also evaluate an agent that can retrieve additional read-only security telemetry during investigation. Access to this evidence does not reliably neutralize the attack, and successful downgrades persist after exposure to poisoned telemetry. Together, these experiments show that telemetry-based IPI can degrade security judgments across model configurations and can persist in tool-using agents with access to corroborating evidence.