Prefix-Tuning for Arbitrary Output Sequences on Pretrained Transformers
Peter Cho-Ho Lam ⋅ Ziyi Wang ⋅ Zirui Zhou
Abstract
Prefix-tuning is a parameter-efficient fine-tuning method for large language models (LLMs) that steers model outputs by prepending task-specific prefix representations to the input sequence. Despite its practical success, the theoretical understanding of the expressiveness of prefixes remains incomplete. In particular, the following question remains open: {\it Given a pretrained autoregressive Transformer such as the GPT series, for any input-output sequence pair $(\mathtt{x},\mathtt{y})$, does there always exist a prefix $\mathcal{S}$ such that prepending $\mathcal{S}$ to $\mathtt{x}$ induces the model to generate the target output $\mathtt{y}$?} In this work, we prove that under the assumption that $\mathtt{y}$ is in the {\it range} of the pretrained Transformer, such a prefix $\mathcal{S}$ exists and can be explicitly computed by solving multiple linear systems. We further show that this assumption is necessary, in the sense that such a prefix may not exist when the assumption fails, thereby providing a complete answer to the above question. Our analysis is constructive and the obtained theoretical results hold under fairly general assumptions on model architecture. As a notable example, the whole GPT-2 architecture satisfies our model assumptions.
Chat is not available.
Successful Page Load