SH$^2$: A Mathematician-Curated benchmark for Assessing Research-level Math Capabilities of LLMs
Abstract
Mathematical reasoning benchmarks for large language models face three persistent limitations. First, datasets scraped from public competitions are vulnerable to training-data contamination and saturate quickly. Second, most existing benchmarks cover only part of the olympiad-to-research spectrum. Third, benchmarks that withhold problems indefinitely to mitigate leakage trade away transparency and reproducibility. We introduce \gls{soohak}, a benchmark of 1{,}141 newly authored problems by 86 mathematicians designed to address all three. \gls{soohak} consists of three difficulty splits, constructed adversarially against panels of small, mid-size, and frontier baseline LLMs, spanning olympiad-style problem solving through research-adjacent material; a separate 99-item Refusal split probes recognition of ill-posed prompts. The dataset is temporarily embargoed and evaluated by request, with a committed public release in late 2026. On \splitthree{}, Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach Avg@3 of 30.39\%, 26.37\%, and 10.39\%, respectively; the best Pass@3 is only 44.12\%, and 129/340 problems remain unsolved by any closed model evaluated, indicating substantial room for improvement before LMs can meaningfully assist in mathematical research. By contrast, the \calibrationsplits{} are nearly saturated by both closed and open-weight systems. Notably, while GPT-OSS-120B, Kimi-2.5, and GLM-5 each reach Pass@3 between 87.4\% and 88.3\% on \splitone{}, within 6 percentage points of the leading closed model (Gemini-3-Pro, 93.85\%), they trail by over 20 percentage points on \splitthree{}. This pattern suggests open-weight pipelines are well targeted to olympiad-style benchmarks but transfer poorly to research-adjacent material. Finally, through a timed human study with 25 participants of diverse mathematical expertise, we confirm that the difficulty is genuinely mathematical rather than artifactual, with aggregated teams covering 50.6\% of a 79-problem sample.