SkillCIR: Intent-Guided Skill Composition for Training-Free Composed Image Retrieval
Abstract
Composed Image Retrieval (CIR) retrieves a target image given a reference image and a textual edit. Existing training-free methods prompt a frozen multimodal large language model (MLLM) to infer the user's edit intent and rewrite the composed query into a single retrieval artifact, such as a target caption, an edited image, or a modality-fused embedding. These methods can reason about the edit, but they have limited means to execute that reasoning during retrieval. A key remaining bottleneck for training-free CIR appears to be not only intent understanding, but intent execution. To address this challenge, we propose SkillCIR, a training-free framework that tackles the intent execution gap by routing the inferred constraints to specialized retrieval skills. Specifically, the same frozen MLLM that infers the intent also emits a structured plan of active constraints; each constraint activates a typed skill that scores gallery candidates along one evidence channel, and the per-skill scores are composed in score space with explicit signs and re-checked by a local verification step. Across three benchmarks and three CLIP backbones, SkillCIR improves the primary recall/mAP metrics by 2.8 to 7.9 points at the latency of typical training-free CIR pipelines and around 23x faster than the latest training-free state of the art. These gains hold with a fixed skill pool and no model-parameter updates, suggesting that improving the means to execute the inferred edit is a useful complement to stronger reasoning about it.