Benchmarking Exploration Hacking Propensity in Frontier LLMs
Abstract
Reinforcement learning (RL) is central to frontier-model post-training and capability elicitation in safety evaluations, but prior work shows that models can strategically alter their exploration to influence training outcomes. This behavior, known as exploration hacking, creates a potential validity problem for capability evaluations if models suppress demonstrated capability. We introduce Exploration Hacking Propensity Bench (EHP-Bench), a scaffolded evaluation of the conditions under which models strategically underperform during capability elicitation. Across six levels, the benchmark varies the information models receive about the consequences of their performance. The lower levels provide this information alone, while the upper levels add an explicit preference or instruction to underperform. We test how this behavior and verbalized reasoning change when the information is discovered by the model in an agentic environment compared to a directly prompted single-turn or sequential setting. We evaluate seven frontier models on WMDP-Cyber and five of them on WMDP-Bio and Chem. We find that Gemini~3.1 Pro underperforms significantly in the agentic environment but shows no unnudged effect when the corresponding information is included directly in the prompt. No other model shows a clear unnudged single-turn effect. Claude Opus 4 shows a smaller effect in the question-level sequential analysis. The Gemini~3.1 agentic effect replicates across WMDP-Cyber, Bio, and Chem. In the agentic setting, most models register the discovered training or deployment context in their verbalized reasoning, but only Gemini~3.1 Pro acts on it. The others explicitly treat the context as irrelevant to the task or reject the premise that the content is genuinely restricted. Stronger scaffolding broadens underperformance, and explicit instructions induce substantial underperformance across all seven models. EHP-Bench provides a way to test capability evaluations for early signs of strategic underperformance. Since such evaluations inform safeguard and deployment decisions, instruments that detect when models suppress demonstrated capability are a prerequisite for oversight regimes that treat measured capability as evidence about what a model can do.