PINT: Effective On-Policy Self-Distillation via Privileged Information Internalization
Abstract
On-policy self-distillation (OPSD) converts training-time privileged context into dense token-level supervision without a separate teacher model: student and teacher share one policy and differ only in what they condition on. This is appealing for long-horizon agentic tasks, where terminal rewards are sparse and credit assignment is hard. But the teacher conditions on privileged information the student never sees at inference, so its token preferences need not be reachable from the base history—a context mismatch that invites shortcut learning and unstable training. We introduce PRIVILEGED INFORMATION INTERNALIZATION (PINT), which internalizes privileged context into parameters rather than leaving it in the prompt, reducing OPSD to a context-matched OPD problem. The current policy proposes trajectories under both base and privileged contexts, and task reward verifies their outcomes; an adapted copy of the policy fits this evidence under the base context alone, yielding a transient teacher that supplies a reward-directed reverse-KL target on the student’s own rollouts for training. We study practical recipes for building such teachers: verified inner objectives, proposal-compatibility weighting, and control of the inner optimization. Over the corresponding strong self-distillation baselines, PINT achieves a 31.6% relative improvement in Qwen3-1.7B WebShop exact success with offline-generated task skills and a 9.0% relative improvement in ALFWorld episode-micro seen-task success with online self-generated skills.