When Learning Signals Persist: Hidden Curricula in Vision–Language Reinforcement Learning
Yujin Jo ⋅ Hyunjik Jo ⋅ Yireun Kim ⋅ Jinsik Lee ⋅ Taesup Kim
Abstract
A fixed task mixture does not guarantee a fixed update distribution in multi-task vision--language reinforcement learning. Under outcome-reward GRPO, only mixed rollout groups containing both successes and failures produce non-zero relative advantage, and queries differ in how long they remain in this state. We analyze these query-level dynamics using the target model's pre-RL solvability and \emph{response demand}, defined as the response length elicited by an external thinking VLM. Among queries with similar initial solvability, higher-demand queries remain mixed longer. Zero-variance filtering converts this difference into accepted-exposure drift, creating a hidden demand curriculum. Its task-level direction is not fixed: reasoning-associated queries remain mixed longer under Qwen2.5-VL Instruct, whereas perception-associated queries do so under Qwen3-VL Thinking. Filtering therefore favors whichever queries the current training setting leaves partially solved. Persistent eligibility, however, need not imply greater training utility. In a controlled comparison matched for initial solvability and category, low-demand training provides broader high-$k$ coverage than high-demand training, including on high-demand evaluation queries. We therefore apply inverse-acceptance calibration, which adjusts candidate sampling without changing the filtering criterion. A past-only EMA controller removes $66.5\%$ of demand-dependent exposure drift and improves high-$k$ coverage with $1.02\times$ the generated tokens of standard filtering, without task labels or future information. Apparent multi-task imbalance can arise before objective conflict, through query-level eligibility and online selection.
Chat is not available.
Successful Page Load