SWE-GPU-Bench: Can Language Models Solve Real-World GPU Software Engineering Tasks?
Abstract
Existing benchmarks for LLM-based GPU programming have made substantial progress in evaluating whether models can generate correct and efficient kernels, operators, or standalone CUDA programs. However, real GPU software engineering often requires modifying existing repositories rather than writing isolated computational units: changes may span CUDA kernels, C++ host logic, Python bindings, build configurations, tests, and performance evaluation scripts. We introduce SWE-GPU-Bench, a benchmark for repository-level GPU software engineering. To the best of our knowledge, SWE-GPU-Bench is the first benchmark that jointly evaluates PR-derived GPU bug fixes, feature implementations, and performance optimizations across multiple real-world repositories and programming languages, while making setup commands, correctness tests, and optional performance evaluation commands explicit components of each instance. SWE-GPU-Bench contains 608 instances from 23 GPU-related repositories and exposes strong cross-file, cross-language, and long-context demands. Through a comprehensive evaluation of representative methods spanning both file-level and repo-level approaches, we show that current LLMs remain far from effectively solving repository-level GPU software engineering tasks.