Incentivizing Medical Vision Capabilities from Large-Scale Multimodal Pre-training
Abstract
Building capable medical vision-language models (VLMs) is challenging due to the scarcity of domain data and the difficulty of eliciting pre-trained medical visual knowledge. We present HuatuoGPT-5, a medical VLM built on a simple yet effective recipe: large-scale multimodal pre-training followed by reinforcement learning (RL). We first construct PubMedVision-Plus, the largest medical image-text dataset with 60M pairs from public PubMed literature, along with a clean 12M high-quality subset. After pre-training on this data, we find that traditional supervised fine-tuning (SFT) struggles to elicit the pre-trained knowledge. In contrast, by applying RL with verifiable and rubric-based instances synthesized directly from the pre-training data, we successfully incentivize the model to use this pre-trained medical visual knowledge. On PubMed-derived data, HuatuoGPT-5 outperforms similarly sized VLMs on multiple medical benchmarks and expert evaluations, significantly narrowing the gap between medical VLMs and leading proprietary VLMs. We hope these data, models, and findings can help accelerate open research on capable medical VLMs.