CONE: A Benchmark for Verifying Counting Reliability in Open-Weight LLMs
Ishaan Verma ⋅ Arsheya Yadav ⋅ Tanvi Bajpai ⋅ Shubh Maheshwari ⋅ Lakshay Verma
Abstract
Verifying whether an LLM has correctly executed even basic numerical reasoning remains brittle: aggregate accuracy hides systematic failure modes, parsers conflate intermediate tokens with final answers, and contradictory premises rarely trigger revision. We present CONE, a diagnostic verification benchmark of 2,500 English question-answer pairs spanning 10 counting-primitive categories, paired with a graded error taxonomy (off-by-one, close, far, non-numeric) and a relative-error protocol that is robust to answer magnitude. We evaluate five open-weight models (gpt-oss-120b, gpt-oss-20b, Qwen3-32B, Llama 4 Scout, and Llama-3.1-8B-Instruct) under matched zero-shot prompting on Groq-hosted inference. The two gpt-oss models lead at $\approx 92.7\%$, Llama-3.1-8B-Instruct trails by $35.44$ pp, and contradiction handling is the hardest category for every model, dominating the universally-failed slice (32 of 52 impossible items despite being only 10% of the dataset). Item-level fragmentation correlation ($n=2{,}500$) and a 116-pair tokenization perturbation study both show that surface tokenization explains $<2\%$ of correctness variance, isolating multi-step character processing and evidence revision as the dominant residual failure modes. A prompt-format robustness check exposes a verification-pipeline failure mode whose mitigation we document. CONE thus serves both as a leaderboard-style benchmark and as a verification-protocol design instrument.
Chat is not available.
Successful Page Load