SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning
Abstract
Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. We propose Self-Verification and Self-Rectification (SVSR), a framework that explicitly integrates verification and rectification into the reasoning pipeline of a compact 7B vision-language model. SVSR follows three stages. First, we construct self-correction trajectories by refining model-generated reasoning with forward and contradiction-based verification. Second, we perform cold-start supervised fine-tuning on 5K high-quality trajectories to initialize structured, multi-step reflective behavior. Third, we apply semi-online Direct Preference Optimization using approximately 20K curated preference pairs, including model-generated candidates filtered by a teacher VLM. Across mathematical and general multimodal benchmarks, SVSR improves reasoning accuracy and generalizes to unseen task types. The trained model also performs better when explicit reflective traces are omitted, suggesting that the learned behavior can support implicit reasoning. These results show that data-efficient post-training can improve the reliability of a compact multimodal model without claiming reductions in inference cost.