What Happens When the Harness Comes and Goes? A Controlled Study of On-Policy Harness Self-Distillation
Abstract
Language-model agents are routinely deployed with inference-time harnesses: procedural templates, strategy documents, and other scaffolding that improves behavior without changing parameters. The benefit is rented: it costs tokens every episode and does not by itself persist when the context is absent. On-policy harness self-distillation moves the harness into the weights: the student acts without it, a teacher frozen at the round's start scores the same tokens with the harness attached, and the student is pulled toward that teacher token by token. We take this recipe as given and measure what persists when the harness is removed and what its restoration adds, in multi-turn settings. The objective never touches environment reward, which suits social reasoning tasks, where outcomes can be a poor training signal but human knowledge of how to act is easy to write down. We study Qwen3.5-27B in database exploration and heads-up poker against an exploitable opponent, with held-out questions, fresh paired deals, temperature-matched sampling, and placebo documents. After one round the distilled model, deployed without the harness, reaches about the level the untrained model reached only with it; placebo distillation does not reproduce this, and in the database domain same-teacher supervised fine-tuning scores lower once the harness is restored. Restoring the harness still helps after one round; after a second round in poker, the harness-free model is near the first round's harness-attached point estimate, a positive harness effect is no longer detected, and the best level observed does not rise. In a small negotiation probe against an LLM counterpart, with no document, a poker-distilled checkpoint acts on every turn and closes every deal while the untrained model hits the generation limit without acting in three: a preliminary sign of transfer beyond the training game. Within its limits (one model family, small evaluation sets, one opponent strategy), the evidence supports a clear conclusion: on-policy harness self-distillation yields persistent, reward-free improvements, and persistence without the harness and responsiveness to its restoration deserve separate measurement.