Mind the Gap: The Divergent Rebound Dynamics of Diffusion and Autoregressive Model
Abstract
Post-training alignment can be fragile: in autoregressive models (ARMs), subsequent fine-tuning on counter-aligned or distribution-shifted data can erode previously aligned behavior. While this rebound phenomenon has been studied mainly in ARMs, how such fine-tuning triggers a similar rebound in diffusion language models (DLMs) remains poorly understood. We conduct extensive fine-tuning experiments across diverse datasets, perturbation sizes, model scales, alignment algorithms, and inference hyperparameters, and find that DLMs consistently exhibit a slower rebound pattern than ARMs do. Furthermore, while ARMs fine-tuned on more positive data suffer a steeper degradation under reverse updates, DLMs exhibit ordering stability: models fine-tuned on larger positive alignment sets retain higher performance as negative data increases. We propose a compression perspective to account for this behavior. Unlike ARMs, whose compression naturally follows a fixed left-to-right order, DLMs generate text through iterative masked denoising. We formalize this distinction through a template-wise compression theory for DLMs. The resulting elasticity theory explains why counter-aligned updates lead to much slower alignment erosion in DLMs. Our experiment findings indicate that rebound is shaped not only by post-training data and objectives, but also by the underlying generation mechanism.