arXiv:cs.AI· Ziquan Zhu, Hanruo Zhu, Si-Yuan Lu, Morris Yu-Chao Huang, Yicheng Lin, Wei Han, Tianlong Chen, Mingyuan Wu, Hanchao Yu, Gaojie Jin, Lu Liu, Bo Sun, Tianjin Huang·· 4 小时前AI 评分32
MOTIVE:面向视觉语言模型的多视角自验证与可靠性引导重思考框架
When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models
AI 导读
研究者提出 MOTIVE,一个面向视觉语言模型的多视角自验证框架,通过正确性锚定的多视角验证学习为每个候选答案打可靠性分数,并在推理时据此做出"接受或重思考"决策。该框架无需外部评判模型,在多种多模态基准和 VLM 主干上持续优于现有自验证与自纠正基线,并减少不必要的推理轮次。
正文
Authors:Ziquan Zhu, Hanruo Zhu, Si-Yuan Lu, Morris Yu-Chao Huang, Yicheng Lin, Wei Han, Tianlong Chen, Mingyuan Wu, Hanchao Yu, Gaojie Jin, Lu Liu, Bo Sun, Tianjin Huang
Abstract:Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verification criterion or fixed prompt, resulting in incomplete and unstable reliability estimates. We first systematically analyze how verifier capability and prompt design affect verification performance. Our findings show that stronger verifiers provide more reliable judgments, while verification performance is highly sensitive to prompt choice, with no single prompt consistently dominating across tasks. Guided by these findings, we propose \texttt{MOTIVE}, a \textbf{M}ulti-View Self-Verificati\textbf{O}n wi\textbf{T}h Rel\textbf{I}ability-Guided Selecti\textbf{VE} Rethinking framework for reliable multimodal reasoning. \texttt{MOTIVE} evaluates each candidate answer from complementary verification perspectives and learns a correctness-aligned reliability score through correctness-grounded multi-view verification learning. During inference, this score governs an accept-or-rethink decision, allowing reliable answers to be returned directly while uncertain ones trigger history-guided rethinking. Extensive experiments across diverse multimodal benchmarks and VLM backbones demonstrate that \texttt{MOTIVE} consistently outperforms strong self-verification and self-correction baselines. Further results show that reliable verification improves accept-or-rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self-verification without an external judge.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07018 [cs.AI] |
| (or arXiv:2610.07018v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07018 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ziquan Zhu [view email]
[v1]
Sun, 4 Oct 2026 15:08:07 UTC (9,818 KB)
来源:arXiv:cs.AI · arxiv.org