Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Abstract
ge language models are becoming a practical interface to optimization: given a natural-language description of an operational problem, they can draft the mathematical model and the solver code, and reinforcement learning with verifiable rewards (RLVR) is the natural training recipe, since a solver can certify every proposed solution. The instances that matter most, however, break this recipe: when a model cannot produce any correct solution, every rollout receives \textit{zero} reward and learning stalls. We measure this learning cliff directly on OR benchmarks: for Qwen2.5-7B-Instruct, roughly half of IndustryOR and OptMATH-Bench (49--57\%) yields no reward in 64 attempts. Privileged guidance, such as prefixes of reference formulations, restores the learning signal but silently changes the question: rollouts are then generated from a training prompt that deployment never sees; we call them \textit{off-context}. We introduce Off-Context GRPO (OC-GRPO), a minimally modified GRPO that keeps guided exploration but applies a change of measure, the importance-sampling device familiar from stochastic simulation, steering the update back to the original unguided objective and provably discounting successes that lean on the guidance. On mathematical reasoning, OC-GRPO achieves a 3.8\% absolute improvement (13.7\% relative gain) over vanilla GRPO with negligible additional cost. On the OR benchmarks, verified reference solutions constructed with stronger models unlock 98.6--100\% of the measured cliff, turning the hardest instances into trainable ones.