The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
Abstract
Continuous image editing requires fine-grained control over semantic attributes in text-conditioned generative models. Recent methods obtain such control through trainable adapters, test-time optimization, or architecture-specific design choices. In this work, we revisit steering the text-encoder space as a simpler alternative for open-vocabulary slider construction. We show that, as text-conditioned generative models become increasingly capable, linear directions in their frozen text-encoder spaces can provide a competitive interface for continuous control when carefully constructed. To make this interface practical, we introduce an automated pipeline that constructs continuous sliders from just a prompt and a concept, eliminating manual data curation and coefficient tuning. First, a language model generates balanced contrastive prompt pairs. Then, difference-of-means over the frozen text encoder estimates a concept direction, and an LLM-assisted process selects the prompt tokens to steer. Finally, an elastic range search adaptively calibrates the slider coefficient interval. As the process includes only changing the text encoder hidden states, the same pipeline can be applied across different text-conditioned generative backbones without denoiser-specific modifications. Across diverse set of edit types, our method outperforms prior training-free baselines and approaches the edit compliance of trained controllers while avoiding per-concept and per-backbone training. Our results suggest that a carefully designed text-space steering is a surprisingly strong practical baseline for open-vocabulary continuous editing. Especially in a rapidly evolving generative model ecosystem, where repeatedly adapting specialized controllers to new backbones can be a substantial burden