STRIVE: Auditing Execution Enabled Reasoning Agents Beyond Accuracy
Abstract
Tool-using language agents are commonly ranked by final-answer accuracy, al- though deployment cost and failure often arise inside the interaction trace: an answer may be numerically correct yet unsupported by executed evidence, reflect- ing parametric recall or an unverified guess rather than evidence-based problem solving; on longer or computation-dependent tasks, this can produce brittle, silent failures. Fluent steps may move away from a solution, and most tokens may be spent before or after decisive evidence appears. Existing holistic model evalua- tions and interactive agent benchmarks expose important capabilities, but do not jointly operationalize correctness, executable provenance, step quality, token util- ity, reasoning-aware redundancy, reliability, and latency for sandboxed reasoning agents. We introduce STRIVE, an audit-first evaluation framework that separates these quantities instead of hiding them in one score. Correctness is resolved by de- terministic equivalence with a selective judge fallback; grounding is a deterministic relation between the declared final answer and successful sandbox output; verified success is their conjunction. A hybrid process metric combines two process reward models (PRMs), execution signals, repetition checks, and a selectively invoked critic. Token utility is anchored to the first traceable solution point, and redundancy penalizes only repetition that neither advances nor verifies the solution. We evaluate six agents on a frozen set of 200 MATH-500 and 100 text-only OlympiadBench problems, yielding 1,800 attempts. The study reveals the correctness–grounding gap and ranking changes that accuracy alone conceals: MiniMax-M3 has the high- est operational correctness (0.760) but only 0.390 verified success, whereas GPT-5 Nano reaches 0.610 verified success; GLM-5.2 is less accurate than MiniMax-M3 but requires the fewest tokens per verified solve (3,323). We release an auditable artifact containing trajectories, code, observations, evidence pointers, evaluator decisions, and analysis tables. The results support multidimensional reporting and Pareto analysis, not a universal scalar leaderboard.