Red-Teaming the Reward: A Verifier and Benchmark for Triton Kernel Optimization Agents
Abstract
Benchmarks for coding agents increasingly reward execution speed, but reported speedups may arise from invalid shortcuts or altered program behavior rather than genuine optimization. This problem is especially acute in GPU kernel optimization, where reduced numerical precision or CUDA Graph replay can produce apparent speedups without improving the requested kernel. We present a verifier for Triton kernel optimization, designed from observed frontier-agent behavior, failures in existing verifiers, and adversarial shortcuts. To evaluate it, we introduce KernelQuest: 100 Triton optimization tasks spanning four levels and 21 task families, each with a certified solution that defines an achievable performance target. Six frontier agents attempted each task three times, yielding 1,800 runs. Our verifier rejected 813 runs before their speedups were credited, including 228 that met their targets. In an independent audit, an expert identified 25% of audited submissions as invalid; three published verifiers and a stopwatch baseline accepted almost all of them, whereas ours accepted almost none. Our verifier accepted all 100 certified solutions. Under adversarial testing, 26 completed runs relied on an exploit, and none was rewarded. Together, these results provide evidence that our verifier is robust to current frontier-agent behaviors, while underscoring that verifier design must evolve as agents become increasingly capable of exploiting reward signals.