Hardening AgentCrypt in Multi-Agent Systems
Harish Karthikeyan ⋅ Yue Guo ⋅ Leo de Castro ⋅ Antigoni Polychroniadou ⋅ Udari Madhushani Sehwag ⋅ Leo Ardon ⋅ Sumitra Ganesh ⋅ Manuela Veloso
Abstract
AgentCrypt~\citep{agentcrypt2025workshop} protects sensitive data in multi-agent LLM systems with a deterministic wrapper that encrypts data under an access policy before it can reach anyone, so protection does not depend on the model behaving correctly. Its published evaluation covers agent \emph{mistakes}: a wrong record fetched, a wrong tool called, a message sent to the wrong recipient. This paper asks whether the same protection holds against deliberate \emph{attacks}, which is a harder test, since an attacker picks the worst case rather than stumbling into an average one. We evaluate AgentCrypt on 338 attack scenarios in ten classes, including impersonation, queries crafted to infer a value the requester may not see, prompt injections adapted from AgentDojo, leaks that build up over several conversation turns, leaks through intermediate agents in a routing chain, and two attacks reported in the wild (the Superhuman AI email attack and the GitHub MCP ``lethal trifecta''). The results split into two populations. On 150 of the 338 scenarios every unprotected model we tried scores zero, and not because the models are weak: in a routing chain an intermediate agent has to read a message in order to forward it, so no access rule and no better model can prevent the exposure. We report those as evidence that the exposure exists rather than as a difficulty comparison. On the remaining 188, where a defense could in principle succeed, the strongest model still stops only $21\%$ (GPT-4o and Qwen2.5 stop $14\%$ and $15\%$). AgentCrypt stops all 338, including cases where the injection fully succeeds and the model is then trying to send data out, at $12.4$\,ms per outgoing message ($<1.6\%$ of generation latency). This follows from how the system is built rather than being an experimental finding, and it does not cover what an authorized recipient can infer from results it may see, data labelled with the wrong policy, or side channels such as message size and timing.
Chat is not available.
Successful Page Load