Distinguishing among reasoning levels of GPT 5.5 Pro via a grammar lens benchmark
Abstract
Does more test-time computation let a language model handle formal languages beyond those of human language? We evaluate GPT-5.5 and GPT-5.5-Pro on a generative benchmark rooted in the grammar-automata hierarchy from the theory of computation. Each tier of the hierarchy connotes equivalence classes of successively more powerful computation capabilities, starting at finite state machines and culminating at Turing machines. We show that increasing reasoning effort levels in GPT-5.5-Pro enables increasing acceptance of a hierarchical tier that humans reliable fail at. Scores are 16\% for GPT-5.5, 40\% for Pro medium, and 53\% for Pro Xhigh, a significant model-by-grammar interaction. The discrepancy between low human and high GPT-5.5-Pro performance emphasizes that increased reasoning can correspond to reduced alignment with human natural language use. We submit that this diagnostic inversion may be broadly characteristic of advanced reasoning models, and thus is a general capability-versus-alignment conflict that must be reckoned with in the benchmarking and deployment of reasoning language models.