Local Distillation: Learning from Monotone Revisions to Student Drafts
Abstract
Preference optimization approaches to improve LLM reasoning usually builds training data pairs (chosen vs rejected) from ground-truth answers or oracle verifiers. Such labeled preference datasets for reasoning tasks are expensive to curate, so recent work on has studied how simply a \textit{positive delta} in preference datasets could be sufficient to train with preference algorithms like DPO. These approaches use a pair of LLMs---one weaker than the other; however, this na\"ive approach does not guarantee positive delta throughout all data pairs. To resolve this, we present a data curation protocol, "monotone revisions." On a prompt, a student presents draft answers to a stronger teacher LLM, which responds with an answer at least as good as the drafts. On a knowledge-intensive Wikipedia reasoning dataset MuSiQue, we find that this protocol creates more positive delta datasets than the na\"ive approach, and training a student LLM improves the multi-hop reasoning performance---we call this technique ``local distillation''. On the flip side, local distillation costs acccuracy relative to the na\"ive approach where the teacher does not see drafts, indicating that conditioning on drafts limits the teacher. We complement these empirical observations on LLM post-training by formalizing the monotone revisions protocol in an idealized contextual linear optimization model and giving efficient, tight algorithms for online learning in this setting. Our theoretical results suggest how our monotone revisions protocol yields preference datasets with sufficient distillation signal, without a need for costly ground-truth answers or perfect oracles.