Vortex: Efficient and Programmable Sparse Attention Serving
Zhuoming Chen ⋅ Xinrui Zhong ⋅ Qilong Feng ⋅ Ranajoy Sadhukhan ⋅ Yang Zhou ⋅ Michael Shieh ⋅ Zhihao Jia ⋅ Beidi Chen
Abstract
Sparse attention is becoming increasingly important for serving large language models (LLMs) as generation lengths continue to grow. However, deploying and evaluating new sparse attention algorithms at scale remains highly engineering-intensive, slowing both human researchers and AI agents in exploring the sparse attention design. To address this challenge, we present Vortex, a system that combines a Python-embedded frontend language atop a page-centric tensor abstraction for expressing a broad range of sparse attention algorithms, with an efficient backend tightly integrated into modern LLM serving stacks. Vortex enables rapid prototyping, deployment, and evaluation of sparse attention algorithms, effectively translating their theoretical efficiency gains into real-world throughput improvements. As a result, Vortex substantially accelerates the design and iteration of sparse attention algorithms. Using Vortex, we further demonstrate AI-agent-driven exploration of the sparse attention, automatically generating and refining diverse algorithms that achieve strong accuracy-throughput trade-offs, with the best generated algorithms delivering up to $3.46\times$ higher throughput than full attention while preserving accuracy.
Chat is not available.
Successful Page Load