Latent-Seed Sensitivity in Prompt Repair: Implications for Small Language Model Agents
Abstract
Prompt-repair agents for text-to-image generation select actions from semantic failure descriptions, although the utility of those actions may depend on latent generation factors that are not observable to the agent. We study this issue in a controlled prompt-repair setting built on Stable Diffusion 3.5 Medium and GenEval2. Our frozen benchmark contains 12 prompts, five latent seeds per prompt, three deterministic repair strategies, and 39 clean count-failure states in which object and attribute constraints remain satisfied. Across valid prompt--repair groups, 58.3% exhibit success flips and 79.2% exhibit sign changes in the continuous count-repair effect across seeds, indicating substantial seed-conditioned variability. We then evaluate Qwen2.5-1.5B-Instruct as a zero-shot small-language-model repair-policy agent using only observable semantic diagnostics while hiding the latent seed. The agent collapses to a single repair strategy for all 39 states and achieves a 10.3% count-repair success rate, below the best fixed strategy at 23.1%, while an oracle reaches 33.3%. Repeated observable-state groups require different oracle actions across latent seeds in 83.3% of cases. These results suggest that single-rollout semantic observations may provide insufficient information for reliable prompt-repair action selection, motivating seed-aware or uncertainty-aware decision mechanisms for small agentic models.