Eviction as a Scheduling Decision in LLM Inference
Xinyu Liu ⋅ Zheng Guo
Abstract
LLM inference faces operational challenges brought by growing KV caches and uncertain output lengths. We formulate serving as a semi-Markov decision process that jointly controls admission, batching, eviction, and re-admission. Batch composition and cache operations determine transition times, capturing the trade-off between computational efficiency, memory use, and eviction overhead. Using processing and transfer times calibrated from Vidur and vLLM, we train a PPO policy and compare it with FCFS scheduling and passive eviction. The learned policy achieves lower latency and higher throughput, showing the costs of active eviction can be outweighed by the improved batching opportunities enabled by released GPU memory.
Chat is not available.
Successful Page Load