SemBench: Exhaustive Semantic Stress Testing for AI Reasoning
Abstract
Reliable evaluation of AI reasoning is difficult when correctness depends on meaning rather than surface form. Static benchmarks can hide failures caused by representation changes, while approximate or human-based judging makes large-scale diagnosis expensive. We introduce SEMBENCH, a fully synthetic benchmark laboratory for exact semantic evaluation of structured reasoning. The benchmark generates finite-domain logical expressions, assigns each expression an executable truth-table semantics, and constructs controlled meaning-preserving and meaning-changing transformations. Because every state in the semantic universe is enumerable, semantic equivalence is decidable exactly and does not require a learned judge, external dataset, model training, or proprietary API. We provide a formal task definition, an exhaustive enumeration protocol, a metamorphic test suite, and metrics separating semantic correctness from surface-form correctness. In an exact enumeration of 1,800 depth-two expressions, only 64 distinct semantic functions occur, demonstrating substantial representational redundancy. Controlled perturbations change semantics at rates of 70.22%, 73.44%, and 63.56% for operator, threshold, and variable mutations, respectively. At depth three, exhaustive enumeration expands to 108,000 syntactic expressions while remaining exactly decidable. These results establish a practical controlled laboratory for future evaluation of language models, reasoning models, agents, and neuro-symbolic systems, while avoiding the confounds of external data and approximate semantic judges.