Token Filtering: Online Attention Pruning via KV Similarity for Efficient LLM Inference
Abstract
Pruning has emerged as a promising direction for accelerating large language model (LLM) inference. However, many existing methods rely on offline calibration data, making them sensitive to distribution shifts between calibration data and inference inputs. In this paper, we introduce Token Filtering, a lightweight online pruning method that selectively skips attention computation for redundant tokens during inference, without requiring any calibration data or fine-tuning. Token Filtering identifies redundant tokens based on joint key–value (KV) similarity and bypasses their attention computation. This approach reduces both compute and KV cache size while preserving essential contextual information. To preserve accuracy while meeting the target pruning ratio, we restrict pruning to later layers, which are typically less sensitive to pruning, with a layer-wise threshold that adaptively tracks the target pruning ratio. Extensive experiments on LLaMA3.1-8B, LLaMA3.1-8B-Instruct, and Qwen3-8B demonstrate that Token Filtering consistently achieves better accuracy–efficiency trade-offs than prior methods. In long-context generation, Token Filtering achieves up to 1.5X higher accuracy than the best-performing pruning baseline while reducing latency by up to 44% compared to the dense baseline.