Anisotropic Trust Region Warping for Prior-Data Fitted Networks in Bayesian Optimisation
Samuel Younger ⋅ Tinkle Chugh ⋅ George De Ath
Abstract
Prior-Data Fitted Networks (PFNs) replace the hyperparameter refitting of Gaussian Process (GP) surrogates in Bayesian optimisation with a single forward pass, at the cost of a prior that is fixed at meta-training time. We show empirically that a PFN's predictive log-likelihood falls increasingly behind that of a refitted GP as the objective function's lengthscale grows beyond the mode of the PFN's meta-training prior, and that this deficit carries over to optimisation performance. To correct the mismatch without retraining the network, we introduce Anisotropic Trust Region (ATR) warping. Per-dimension lengthscales are estimated by kernel target alignment, and an affine rescaling of the input space maps the objective onto the lengthscale the PFN expects, with a trust region around the anchor determining which observations are retained, leaving the PFN's $\mathcal{O}(d \cdot N^2)$ inference cost unchanged. On functions drawn from a Matérn-3/2 kernel, ATR warping improves optimisation performance of the PFN at both extremes of the lengthscale range.
Chat is not available.
Successful Page Load