Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning for T2I Generation in Unified Multimodal Models
Abstract
Test-time scaling (TTS) improves text-to-image (T2I) generation in unified multimodal models (UMMs), but existing TTS methods typically rely on frozen-parameter search or sampling, yielding only ephemeral, instance-level improvements. We propose Meta-TTRL, a metacognitive test-time reinforcement learning (TTRL) framework that converts test-time generation experience into learning signals, enabling UMMs to self-improve during T2I generation without external reward models. Meta-TTRL formulates test-time learning as a closed-loop interaction between an object-level generator and a meta-level introspector. The introspector decomposes prompts into structured verification rubrics and produces confidence-enhanced intrinsic monitoring signals, which are used to optimize the generator policy through reinforcement learning (RL). Extensive experiments demonstrate that Meta-TTRL generalizes well across three representative UMMs, including Janus-Pro-7B, BAGEL, and Qwen-Image, achieving significant gains on compositional reasoning tasks and multiple T2I benchmarks with limited data. Further analyses with external introspectors, RL leakage, and alternative monitoring signals reveal a key principle for effective TTRL: metacognitive synergy, where monitoring signals must align with the model's own optimization regime.