Which Prompt Mutations Can an Agent Actually See? Coverage, Mutation, and Static Conflict Detection over the Agent Policy Layer
Abstract
An agent's behaviour is governed by a prompt: a document of rules, conditions, and priorities that is edited like code but has no testing discipline around it. We ask what can be verified about that policy layer deterministically, before any model is called, and what cannot. We define rule, condition/branch and rule-pair coverage over structured prompts, six mutation operators, and a static analysis for rules a higher-priority sibling shadows, proved sound over a guard alphabet including loop frames. The main result is a negative one, found by auditing our own tool. Rule annotations are comments and are stripped before the model sees the prompt, so an operator editing only an annotation cannot change the rendered text and cannot be killed by any rendered-input oracle: it is equivalent with respect to model behaviour however much structural bookkeeping it disturbs. This affected five of our six operators, leaving 81% of mutant kills invisible in the rendered prompt and a 5.4 times gap between the two oracles. Repairing them, by carrying instruction bodies with their annotations and refusing to swap rules across control-flow boundaries, leaves three of six model-visible, the invisible share at 39% and the gap at 1.65 times . Both states are measured, each from a pinned commit of the tool. The same 39% appears on a hand-annotated corpus of real agent prompts with a very different shape, consistent with annotation stripping being the dominant source of the gap. A behavioural probe on 80 model-visible mutants finds a small, operator-dependent effect: clear for DropRule, suggestive for NegateCondition, undetectable for SwapRules. On 60 renderable prompts, coverage-directed generation reaches mean rule coverage 0.94 and recognised branch coverage 0.78, and raises mutation score from 0.63 to 0.73 under the structural oracle and from 0.38 to 0.45 under the rendered-input oracle. Two limits we state plainly. Hand-annotating 12 of 40 real-world agent skill prompts, byte-exact to the originals, shows that real prompts are rule-dense but almost branch-free: 455 rules against 9 template-expressible conditionals and no loops. The rule abstraction transfers; the branch criterion has almost nothing to measure. And checking the soundness theorem against the implementation found two false positives, the second of which makes the theorem false over the language the renderer accepted; we repair it by enforcing the missing precondition rather than by weakening the claim.