The Price of Choosing Examples: Context Overfitting in Adaptive In-Context Learning
Xiaoyu Li ⋅ Jiangxuan Long
Abstract
In-context learning is usually analyzed as if the examples in the prompt were sampled before the learner is chosen. In deployed systems they are not: examples are retrieved, filtered, re-ranked, edited, and repeatedly tested after seeing the query and the data store. We formalize this gap by treating prompt construction as adaptive data analysis through a conditional stochastic-kernel calculus. Our main quantity is the \emph{context capacity}, the mutual information or approximate max-information between the data store and the final prompt transcript. For any frozen in-context learner and any adaptive context pipeline with capacity $\kappa$, we prove a finite-sample context generalization bound of order $\sqrt{\kappa/k}$, where $k$ is the number of examples that enter the prompt. We strengthen it to a sub-gamma transport inequality with Bernstein curvature, prove a matching lower bound, and derive two PAC-Bayesian retrieval rules: a first-order Gibbs rule and a self-normalized second-order rule. Finally, we instantiate the theory for Bayesian linear in-context regression, including an order-statistics analysis of query-aligned top-$k$ Gaussian retrieval. The results separate context quality from context certification: longer or more relevant contexts improve risk only when the procedure used to choose them is stable enough for selected-context evidence to transfer.
Chat is not available.
Successful Page Load