Trajectory-Matching Meta Pseudo-Labeling for Semi-Supervised Learning
Abstract
Meta pseudo-labeling improves semi-supervised learning by updating a teacher according to whether its pseudo-labels help a student improve on labeled data. Existing one-step meta-gradient methods, however, differentiate through the student update, introducing mixed higher-order derivatives, extra memory cost, and noisy minibatch meta-gradients. We propose Trajectory-Matching Meta Pseudo-Labeling (TMPL), a first-order alternative that extracts labeled-data feedback from recent student optimization trajectories. We show that the one-step meta objective is locally equivalent, up to second-order terms in the student step size, to maximizing alignment between the outer gradient and the pseudo-label-induced student gradient. TMPL approximates this alignment without differentiating through the student update by matching finite differences of the outer objective and the pseudo-label-induced student objective along recent student displacements. Its implementation stores only a FIFO cache of detached scalar objective values, rather than student computation graphs or past student networks. We prove that small trajectory-matching error controls the gradient mismatch within the span of recent displacements, with an explicit residual for unexplored directions. On CIFAR10, CIFAR100, SVHN, and STL10, TMPL improves over MPL in all eight evaluated label regimes, obtains the best result in five settings, and remains competitive with strong SSL baselines. On CIFAR100 with 2500 labels, TMPL reduces peak memory by 39.7\% relative to MPL and 54.0\% relative to a raw differentiable meta-gradient implementation. Ablations identify the directional trajectory-matching term as the main source of improvement, while alignment diagnostics show that TMPL recovers MPL-like gradient alignment after a short warm-up.