Selective Answering for Medical VQA via Parallel Independent Claim Verification
Abstract
Medical visual question answering is a safety-critical decision problem in which abstaining on inputs the image cannot support matters as much as answer accuracy. Existing inference-time agentic methods address hallucination by adding a verification step, but they collapse the final decision at a single language-model judge that integrates all evidence in one prompt and is systematically biased toward declaring inputs answerable. We propose PICV (Parallel Independent Claim Verification), which reformulates selective medical VQA as independent claim verification with transparent aggregation: the question is decomposed into a small set of typed visual claims, each is verified by an isolated prompt with no access to the other verdicts, and the verdicts are combined by a deterministic rule rather than another model, yielding an answer/abstain decision together with a typed unanswerability reason. We evaluate PICV efficiently on an automated answerability benchmark we construct over two public radiology datasets via three reproducible corruption procedures, requiring neither heavy VLM judging nor human annotation at evaluation time. Across four VLM backbones, PICV consistently outperforms strong prompting and agentic baselines; ablations attribute the gains specifically to per-claim isolation and rule-based aggregation rather than to agentic decomposition alone.