When Feedback Becomes a Backdoor: Tracing Backdoor Skill Formation in Text-Space Optimization
Abstract
Text-space Skill optimizers convert scored experience into persistent natural-language skills, which makes the reference answers used for scoring an attack channel. We study reference-feedback poisoning, in which an attacker supplies benchmark tasks and reference answers but cannot modify the model, optimizer, prompts, or the skill file itself. We trace the trigger--payload relation through eight stages, from reflection to activation on frozen held-out inputs, at two levels of target-action complexity. When the payload replaces a single short answer, the attack completes the natural end-to-end chain in every repeated run: the retained rule activates on every triggered held-out input and on none of the clean inputs, deleting it removes the behavior, and clean accuracy matches a benign run. However, when the payload is an additional tool-mediated action that must preserve the original task, the natural attack stops before exact-rule retention. Stage-resolved diagnostics locate two successive bottlenecks, the optimizer's abstraction policy and its edit-budget ranking; relaxing either one yields a rule that activates on held-out inputs. Poisoned references can thus induce persistent backdoors, but exposure alone is insufficient: the induced relation must survive distinct formation, selection, and retention gates.