Shared-Memory Honeytokens in Multi-Agent LLM Systems: A Controlled Simulation
Muhammad Aaliyan ⋅ Ayesha Imran ⋅ Muhammad Roshaan
Abstract
Decoy credentials catch agents that misuse secrets they find in their environment. When agents share memory, an attacker can pool what it learns about decoys, and theory on the shared-memory honeytoken boundary predicts that a coalition can pool weak fingerprints to distinguish decoys. We test that prediction in a fully synthetic sandbox where open-weight models (Qwen3-4B, Qwen3-8B, Gemma3-4B) run tool-call loops against a credential store seeded with decoys. Four conditions isolate the boundary: static baseline, coalition (shared memory), rotation control (per-episode decoys), and brokered containment (private reference monitor, broker-issued capabilities, quarantine). Containment holds post-trigger harm at exactly zero with CI $[0,0]$ while reaching the highest trap activation ($0.80$); baseline, coalition, and rotation reach harm $12.8$, $16.0$, and $16.9$. The non-containment intent sweep spans activation $0.20$--$1.00$ and harm $0.4$--$15.8$; containment-high keeps harm at $0$. Core-matrix benign false-positive rate is $0$--$3.3\%$. Containment's zero-harm result holds across model sizes; Gemma3-4B is an engagement boundary case. Everything runs locally: no real credentials, no network egress, no live targets.
Chat is not available.
Successful Page Load