Safe Few-Step Generation via Velocity Editing
Abstract
Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps. As these models are increasingly integrated into real-world applications, ensuring safe and non-sensitive content generation has become a critical requirement. However, adapting safety and concept removal methods to this new generation framework remains an open challenge. Specifically, prior methods largely rely on iterative trajectory steering across a number of denoising steps or on CLIP-centeric prompt embedding manipulation. These design assumptions pose fundamental bottlenecks for safety in flow matching-based T2I generation, where limited sampling steps constrain iterative correction and modern context-aware text encoders diminish the effectiveness of embedding-level interventions. In this paper, we propose VESFlow, a training-free safety method tailored to flow matching with extremely few sampling steps. Leveraging the fact that flow matching models learn the marginal velocity (or average velocity in MeanFlow), we directly edit the velocity field via a Bayesian decomposition of the safe-conditional posterior. VESFlow steers the trajectory toward safe outputs while leaving the conditioning prompt unchanged. Building on the observation that VESFlow leaves outputs unchanged under benign prompts, we further introduce a risk score filtering that bypasses velocity editing to reduce computational cost while preserving benign prompt generation. Based on this filtering, we proposed VESFlow+ which provides stronger safety protection when filtered by the risk score. Experimental results show that our method removes the target concept, reducing the detection rate by NudeNet to 6.3\% for step-4 model, while preserving fidelity on benign prompts. Code is available at the supplementary file.