Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation
Abstract
Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings. For example, prior studies show that cardiomegaly may be omitted from ICU chest X-ray reports when the imaging request is focused on monitoring support device placement. As a result, models trained with standard approaches inherit these omissions, learning to under-report findings themselves. We propose PU-DPO, a preference optimization framework that constructs contrastive pairs via editing model responses, producing variants that explicitly mention or omit a specific finding. To prevent omission noise from corrupting the preference signal, we reformulate the objective under a positive-unlabeled (PU) learning framework, treating absent mentions as unlabeled rather than truly negative. Across semi-synthetic experiments and preliminary analyses on real-world chest radiograph benchmarks where adjudicated labels are available, PU-DPO yields consistent gains in detection rates across multiple pathologies without compromising specificity or overall report quality, and is more robust to omission noise than prior approaches.