PreFT: Prefill-only finetuning for inference efficiency
Andrew Lanpouthakoun ⋅ Aryaman Arora ⋅ Zhengxuan Wu ⋅ Dhruv Pai ⋅ Benjamin Keigwin ⋅ Dan Jurafsky ⋅ Chris Potts
Abstract
Large language models can now be personalised efficiently at scale using parameter efficient fine-tuning methods (PEFTs), but \emph{serving} user-specific PEFTs harms throughput, even with specialised kernels and memory management techniques. This is because, theoretically and empirically, a mismatch exists between prefill (processing a large number of tokens at once) and decode (generating a single token autoregressively): the latter has far lower throughput when serving multiple adapters. Rather than optimising performance relative to parameter count, for efficient multi-adapter serving, we instead ought to optimise performance relative to \textit{serving throughput}. We therefore propose \textbf{\texttt{PreFT}} (Prefill-only Finetuning), wherein we only apply the adapter to prefill tokens and discard it afterwards. \texttt{PreFT} significantly increases throughput with minimal effect on performance. We develop and release an efficient implementation of two prefill-only PEFTs, LoRA and ReFT, on the vLLM inference engine. We first show that serving multi-user \texttt{PreFT}s is vastly more efficient than traditional PEFTs ($1.90\times$ the throughput when serving $512$ adapters on Llama 3.1 70B). Then, we compare the performance of prefill-only vs.~all-token adapters on a variety of supervised finetuning and reinforcement learning tasks with LMs at varying scales. On SFT, we observe that the evaluation loss of \texttt{PreFT}s is higher, but can be compensated by increasing rank with nearly no reduction in throughput. On RL, we consistently find that \texttt{PreFT}s approach parity with standard PEFTs. Together, this work validates prefill-only adaptation of LLMs as a more favourable accuracy--throughput tradeoff than existing PEFTs for personalised serving.
Chat is not available.
Successful Page Load