SARA: Step-Adaptive Rank Adjustment for Diffusion Inference Acceleration
Abstract
Diffusion models achieve strong performance across image and audio synthesis, but their inference remains computationally expensive, requiring tens to hundreds of sequential evaluations of a large denoising network. We observe that the effective rank of the denoiser output varies sharply across timesteps and architectures, motivating step- and block-wise adaptation of computation. We propose Step-Adaptive Rank Adjustment (SARA), a training-free framework that treats the SVD rank of pretrained linear layers as an inference-time control variable, adapted per step and block by SARA-DP and per step by SARA-Online. SARA introduces the Low-Rank Approximation Error (LRAE) — the discrepancy a rank-r truncation induces against the full-rank network — and instantiates it in two algorithms: SARA-DP measures LRAE on a few warmup prompts and solves the globally optimal rank assignment under a MACs budget by dynamic programming; SARA-Online computes a closed-form, spectrum-based LRAE from the model output at inference time, requiring no calibration and no trajectory lookahead. SARA applies uniformly to UNet (SDXL), MMDiT (SD3), DiT (Stable Audio Open), and hybrid MMDiT/DiT (TangoFlux) backbones, spanning ϵ, v, and rectified-flow parameterizations. SARA combines with existing solver- and cache-axis acceleration to yield compounded wall-clock speedups in our experiments. Together, these results introduce rank as a new axis of diffusion inference acceleration.