TritonTune: LLM-Guided Multi-Agent Optimization of GPU Kernel Configurations
Abstract
Tuning Triton GPU kernel launch configurations is a critical but labor-intensive step in production ML system deployment. Existing approaches rely on exhaustive autotuning, which requires manually defined search spaces and becomes prohibitively slow for the long tail of operator shapes. We present TritonTune, a multi-agent system that combines LLM-guided search with closed-loop empirical validation to optimize Triton kernel configurations safely and efficiently. TritonTune pairs two domain-specialized LLM agents with complementary expertise in data movement and compute resource allocation, resolves their proposals through a lightweight orchestration layer, and refines the best candidate through a phased benchmark search on real hardware. Per-agent feedback drives targeted corrections when regressions occur, and a hardware-validated fallback provides a worst-case guarantee absent from prior LLM-based approaches: the system can only improve performance or leave it unchanged. On 19 TritonBench kernels and 393 input shapes spanning hand-written, torch.compile-generated, and quantized operators, TritonTune achieves a 2.24× per-shape geometric-mean speedup on an NVIDIA V100 with 80% of shapes improved, under a closed-loop fallback that guarantees no accepted shape regresses below the original. Cross-GPU runs on L40S, A40, and A100 further confirm that our design generalizes beyond V100. These results show that LLM-guided multi-agent optimization, when grounded in hardware measurement, can serve as a practical, safe complement to exhaustive autotuning.