Thinking with Imitation: Adaptive Reinforcement-Imitation Learning for Tool-Augmented Scientific Reasoning
Abstract
Expert scientists rarely solve genuinely novel problems through parametric recall alone; instead, they reason by analogy to structurally related prior cases and adapt their solution procedures to new contexts. Yet current multimodal large language models (MLLMs), even when equipped with retrieval, computation, and visual tools, still struggle with this imitation-driven form of reasoning. We further observe a counterintuitive negative result: simply prepending expert problems and full solution trajectories as context does not improve performance on novel scientific tasks, and can even slightly degrade it. Motivated by this observation, we argue that scientific reasoning should shift from passive answer conditioning to an imitation-driven paradigm, in which models actively seek relevant precedents and reuse their solution structures. We instantiate this idea with MimicAgent, a tool-augmented scientific reasoning agent that retrieves structurally related exemplars, interprets their reasoning trajectories, and transfers their solution patterns to the target problem through multi-step reasoning and tool interaction. To train this behavior under sparse rewards, we introduce ARIS (Adaptive Reinforcement-Imitation Switching), which unifies on-policy reinforcement learning with online expert imitation. When all rollouts for a prompt fail, ARIS replaces degenerate reinforcement updates with expert-generated demonstrations, yielding an implicit imitation-to-reinforcement curriculum without requiring a cold-start stage or fixed-ratio offline mixing. To support training and evaluation, we further introduce SciExplore-Bench, a bilingual multimodal benchmark for open-ended experimental scientific reasoning, paired with a manually curated repository of analogical exemplars. The substantial gains and impressive generalization across multiple benchmarks suggest that imitation serves both as a global paradigm for exemplar-guided reasoning and as a local mechanism for stabilizing reinforcement learning.