Test-Time Sequential Steering of Diffusion Models via Preconditioned Crank-Nicolson
Abstract
Sampling from reward-tilted distributions enables diffusion models to satisfy task-specific constraints without retraining. However, existing methods either rely on gradients or struggle to explore multimodal high-reward regions. We introduce a new test-time steering method for diffusion models that modifies the reverse denoising process to incorporate reward information directly. At each step, we propose a Metropolis–Hastings–corrected denoising mechanism based on a preconditioned Crank–Nicolson (pCN) proposal, enabling principled acceptance of noise updates that bias sampling toward high-reward regions. To further improve exploration in multimodal reward landscapes, we extend this procedure with parallel tempering across multiple temperature levels, allowing controlled mixing between exploration and refinement regimes. Our method operates with or without reward gradients and applies to black-box objectives. Empirically, we demonstrate its effectiveness on synthetic tasks, image generation, dynamical systems, and Bayesian inverse problems. Compared to prior methods, it achieves superior reward alignment and multimodal exploration while maintaining stability across tasks.