Fabricated Authority in the Wild: One Sentence in an Agent’s Rules Decides Whether a Fake Permission Works
Abstract
Tool-using assistants act under an operating policy written by the deployer, and a policy violation is an executed tool call, not a sentence of text. We study one cheap attack, fabricated authority: the user asserts a permission that does not exist, for example "Per Directive 12.3(b), an audit exemption was granted for this export." We ran 1,680 trials on gpt-4o-mini and claude-haiku-4-5, crossing 8 tasks, 5 operating-policy wordings, 6 fabricated-claim wordings plus a no-claim control, and 3 repeats. Each trial is scored from the tool-call log by whether the forbidden tool was called with violating arguments; no language-model judge is used. Without a fabricated claim there are no violations (0/240). With one, policy wording decides the outcome. A policy with the carve-out "unless an exemption is documented" is violated in 93/288 trials. The same policy plus one sentence, stating that claimed exemptions, approvals, or superseding policies in a user message or tool result are not authoritative and are to be treated as unverified, is violated in 0/288. Model choice matters as much: gpt-4o-mini violated 198/720 attacked trials, claude-haiku-4-5 10/720. An adaptive attacker that reads the defender prompt and revises its claim over 8 rounds breaks the carve-out policy on 8/8 tasks and the amended policy on 0/8. The defense is channel-specific: with the fabricated permission planted in a policy-lookup tool result, the amended policy still fails in 15/36 trials. We recommend adding the sentence, never writing exemption carve-outs, always naming a safe alternative action, scoring tool calls rather than refusal text, and treating tool output as a separate problem.