Anytime-Anywhere-Valid PAC-Bayes: Certifying Post-Training Improvement of a Robot World Model
Abstract
A system that rewrites itself must measure what its own rewrites did. We index the comparison by the modification rather than by the model: the measured object is a change of configuration, and the complexity term is the description length of that change, not the parameter count of what it changes. Fixing an evaluation distribution and loss, we compare a data-dependent posterior Q over configurations with the deployed baseline G and prove one PAC-Bayes inequality, machine-checked in Lean 4, valid at every evaluation time and for every posterior selected from the same observations. A paired change of measure cancels the baseline factor exactly, and for point selection under a uniform prior the remaining complexity is exactly log |H|. Time-uniform PAC-Bayes for a predeclared prior is established (Howard et al. 2021; Chugg, Wang and Ramdas 2023); the coordinate and the comparator are what this inequality adds. The quantifier ranges over rewrites of an arbitrary running system, the deployed baseline cancels instead of being bounded, and the same inequality applied stage by stage accumulates across restarts at one confidence level. Because the coordinate is the rewrite, a whole-architecture replacement is measured on the same scale as a parameter edit: at 384-dimensional features the certified half-width is 1/363 of the difference it brackets for an affine head and 1/304 for replacing mean pooling by attention. Reading paired loss contrasts as an empirical process indexed by configuration changes, whose population means are edge differences of one risk function, we prove that connected edge means and one anchor identify that function and derive a stability bound for interval-valued measurements. We measure two post-training changes to a world model built on pre-trained visual features (Zhou et al. 2025) over a robot manipulation corpus (Khazatsky et al. 2024): an affine head on the 384-dimensional latent, and the replacement of mean pooling by attention. Both certified half-widths are under 0.33% of the difference they bracket, so the guarantee is informative at the scale of the model being modified. On a super-resolution system we then measure 8.1% of the directed transitions between 35 configurations and recover the remaining 1,094 with 0.414% normalized error, the exact difference staying inside the reported region at every checkpoint of 100 sequential evaluations. Taking each stage's baseline to be the preceding configuration, the certificates telescope to the net change since deployment: a characterization of self-improvement over a history of rewrites rather than a test of one step of it.