Salvation Lies Within: Proactive Prefix Re-forming for LLM-based Tagging
SHULAN WANG ⋅ YingJie Zhu ⋅ Yuting Yan ⋅ Ke Cheng ⋅ Yan Chen ⋅ Sheng Zhang ⋅ Hefei Guo ⋅ Qidong Zhang ⋅ Zhengyong Zhang ⋅ Yibo Jin ⋅ Zhuzhong Qian
Abstract
Industrial personalized recommendation increasingly adopts LLMs for offline intent tagging (e.g., user interest from installed apps or browsing history), where each user corresponds to one prompt with thousands tokens. While shared prefix among prompts can improve the inference efficiency, the reusable part is actually dwarfed by user-specific records. We observe that even a slight $\textit{discordance}$ in preceding records prevents two prompts from sharing a longer prefix, revealing an opportunity to proactively re-form the common part from within, by adjusting the internal orders, with ensured accuracy for intent tagging. Realizing this at industrial scale is challenging: (1) the data preparation and the inference operate separately; (2) re-forming must complete within minutes for massive prompts under strict SLO; and (3) user data evolves continuously, requiring incremental updates. We present $\textit{PrefixShakeup}$, a system that enables $\textit{proactive prefix re-forming}$ before actual inference, intrinsically enlarging the common part from the source. PrefixShakeup, bridges the data and inference by enhancing the datasets during its preparation, thereby increasing the reuse. It combines three techniques: (1) a prefix model that tolerates a bounded number of records, instead of requiring strict prefix matching; (2) an alleviator that reorders both records within each prompt and the prompt positions to maximize the reuse under limited HBM, upon the kernel-based acceleration for scoring the similarity; and (3) an adaptor that supports incremental changes via lightweight management and necessary cache over time, to avoid full re-computation. We implement PrefixShakeup, on Ascend NPUs and evaluate it in the mirror environment. Compared to state-of-the-art inference, PrefixShakeup, improves throughput by up to 2.1$\times$, with controlled tagging accuracy.
Chat is not available.
Successful Page Load