Streaming Attention Approximation with Page-Aligned Memory Management
Jane Lee ⋅ Neharika Jali ⋅ Insu Han ⋅ Amir Zandieh ⋅ Jieru Mei ⋅ Jinoo Baek
Abstract
We propose a method for approximating causal attention in a streaming setting under additional memory constraints reflecting practical memory management in inference and serving systems like vLLM. We improve over prior work in providing 1) provable theoretical bounds for a block eviction method, 2) implementation compatible with vLLM, 3) special cases where the method can match token-level eviction methods and 4) initial experiments competitive with other block-level eviction methods and some system-level improvements over the baseline.
Chat is not available.
Successful Page Load