Probe Guided Curriculum Learning: Models Secretly Know Good Prompts from Bad
William Bankes ⋅ Yi Y Ng ⋅ William Gitta Lugoloobi ⋅ Thomas Foster ⋅ Sangwoong Yoon ⋅ Ilija Bogunovic
Abstract
Curriculum Learning seeks to address a common failure in GRPO-based post-training: prompts with zero advantages. These strategies typically use rejection sampling to build informative batches at the cost of sampling additional rollouts. We propose to use success probes, which leverage the internal belief of how likely a model is to succeed, to scan prompts and only roll out those most likely to provide training signal, whilst reducing total rollouts. In math reasoning with a small LLM ($\sim 1B$ params), success probe guided sampling matches the performance of curriculum learning algorithms whilst \emph{using half as many rollouts}.
Chat is not available.
Successful Page Load