Characterizing Inference-Time Adaptation in Medical QA: From Retrieval Augmentation to Reward-Guided Alignment
Yaxuan Wu ⋅ Bangxu Tian ⋅ Yifan Wang
Abstract
In high-stakes medical question answering, a deployable system must not only select correct answers but also express calibrated confidence, avoiding predictions that are confident yet unsupported by evidence. We study how to improve both together, combining two interventions in a single controlled setting: policy optimisation with direct preference optimisation (DPO) over reward-guided preference pairs, and inference-time answer selection that grounds each candidate reasoning trace in retrieved evidence and scores it with a process reward model (PRM). Preference data is built from MedMCQA, MedQA, and PubMedQA, and every configuration is evaluated on held-out MedQA using accuracy together with expected calibration error (ECE), Brier score, and an overconfidence gap. DPO alone improves accuracy but leaves the model overconfident; retrieval-grounded, PRM-verified selection is what suppresses confident but unsupported answers, and combining the two is both the most accurate configuration and, on Brier score and overconfidence gap, the best calibrated, reaching $0.787$ accuracy while cutting the overconfidence gap from $0.329$ to $0.020$ on an 8B backbone. Reward-guided alignment and evidence-grounded verification are therefore complementary routes toward medical QA that is accurate and trustworthy. Our code and data are available at \href{https://anonymous.4open.science/r/inference-time-adaptation-56CC/}{https://anonymous.4open.science/r/inference-time-adaptation-56CC/}.
Chat is not available.
Successful Page Load