Stress-Testing RL-based Capability Elicitation with Randomized Underperformance
Abstract
Reinforcement learning (RL) is increasingly used as an elicitation technique to assess the upper limits of language-model performance, including potentially hazardous capabilities. We study randomized underperformance, an RL-elicitation resistance strategy in which a model follows a randomized in-context procedure for underperforming. Such strategies can yield intermittent rewards and nonzero within-group advantages, while successful rollouts may arise from randomization rather than from the target capability. We instantiate randomized underperformance in prompt-locked models whose underperformance is activated by a trigger and test whether GRPO can elicit the locked capability. On MMLU-Redux, models using simple strategies such as choosing a random letter or deliberately choosing a wrong answer rapidly recover to above 70% held-out accuracy on gpt-oss-120b within 5-10 RL steps. In contrast, models using more structured multi-step modulo procedures remain close to chance after 100 steps despite receiving intermittent reward. We further observe that reducing the trigger frequency systematically increases the number of RL steps required for elicitation. Finally, on open-ended code generation, RL recovers much of the suppressed performance without eliminating the randomization behavior: the model continues to execute the procedure but learns to circumvent its intended effect on the final output. Randomized underperformance therefore provides a concrete stress test for RL-based capability elicitation pipelines.