Learn in 2048 Dimensions, Certify Along One: PAC–Bayes Bounds for Few-Shot Prompt Tuning
Abstract
Few-shot soft-prompt tuning updates 2048 parameters of a frozen vision–language model, and certifying the result is hard not because the fit is poor but because a prior, a posterior, the search over both and a numerical evaluation of the risk must all be stated—and charged for—in that same space. The accounting, not the fitting, is what fails to scale. TaskLine-PB certifies not the fitted prompt itself but a choice between it and zero-shot CLIP. A fitting split produces the full 2048-dimensional update, kept intact rather than compressed or projected; an independent calibration split then fixes only the odds with which a stochastic predictor draws the fitted prompt rather than the pretrained one, and it is that predictor's class-balanced risk that carries the guarantee. Relative entropy is unchanged by the map from the coordinate to the prompt, so the bound pays for the prompt that is deployed and not for a proxy; the best certificate over every posterior on the line has a closed form, which at the two named ends reduces to a formula in two error counts; and the adaptation as a whole costs a single bit, with nothing left to estimate: no Monte-Carlo term, no quadrature, no free parameter. However badly the fitting split is corrupted, the certificate cannot rise above the branch that never looks at it—a theorem rather than an observation. Sixteen labelled examples per class put every certificate on eight CLIP tasks below random guessing, at a cost of 0.0066 on average and at most 0.0144 against a certificate of the fitted prompt alone; to our knowledge this is the first non-vacuous PAC–Bayes risk certificate for few-shot soft-prompt adaptation of a frozen vision–language model without auxiliary-task checkpoints, an upstream task collection, or a learned generator. Nothing about the adaptation was made smaller; only the question the certificate has to answer was.