PTXBench: Can LLM Kernel Agents Directly Optimize PTX on Recent GPUs?
Abstract
Can LLM kernel agents directly optimize kernels through architecture-specific PTX on new GPUs? Existing benchmarks evaluate generated kernels primarily through correctness and latency, but do not systematically study the programming interface through which agents express hardware-specific optimizations. We present PTXBench, a framework for evaluating direct CUDA--PTX kernel optimization from controlled architecture contracts. Models generate CUDA kernels with inline PTX; PTXBench evaluates correctness, confirms use of the requested PTX by observing corresponding SASS instructions execute at runtime, and measures optimization quality as speedup over cuBLAS and cuDNN. Across five models, five GEMM and attention workloads, and H100 and B200 GPUs, the gap between correctness and target-PTX execution is smaller on Hopper, whereas many correct Blackwell kernels do not use the requested PTX; confirmed use on either architecture often remains uncompetitive. Under matched refinement, Triton achieves a higher best speedup in every Blackwell setting and all but one Hopper setting, and generally uses fewer reasoning tokens. Our results expose the new-GPU optimization gap: today's agents can sometimes execute the requested low-level path, but cannot yet reliably turn architecture-specific PTX into competitive kernels.