WirelessMathBench-XL: A Contamination-Audited Benchmark for Wireless Mathematical Reasoning
Xin Li ⋅ Mengbing LIU ⋅ Yiyang Zhu ⋅ WENHE ZHANG ⋅ LI WEI ⋅ Jiancheng An ⋅ Chau Yuen
Abstract
**WirelessMathBench-XL** is a $4{,}027$-problem benchmark for wireless mathematical reasoning, built from $836$ retained arXiv papers across $20$ wireless subfields. Benchmarks built from arXiv papers risk overlap with the same arXiv text used in LLM pretraining, yet few releases include the provenance and metadata needed to audit this overlap. WirelessMathBench-XL builds contamination auditing into the released artifact through a **reverse-probe $13$-gram audit**: it indexes benchmark prompts, streams public pretraining corpora, and records per-problem prompt-surface lexical overlap with benchmark-scale memory. Against $12.6$\,B streamed $13$-grams from RedPajama-arXiv, the audit identifies a strict zero-hit subset $\mathcal{S}_0$ covering $3{,}853$ problems ($95.7\%$). Filtering to $\mathcal{S}_0$ changes accuracy by less than $1$\,pp for every evaluated model; frontier-model calibration rows form a single high-accuracy cluster between $86.5\%$ and $91.3\%$, not a resolved rank order. Thus, on this fixed prompt-surface lexical audit channel, reported scores are not measurably driven by exact recall of detected RedPajama-arXiv prompt text. The audit does not cover paraphrase, target-answer, post-training, or closed-corpus exposure. The release includes source-paper identifiers, verifier-facing ground truths, contamination metadata, cleaned subsets, a paper-disjoint sensitivity view, evaluation traces, paired-bootstrap scripts, training recipes, Croissant metadata, and a Datasheet for Datasets, so users can rerun, tighten, or replace the cleaned view.
Chat is not available.
Successful Page Load