跳到正文
arXiv:cs.AI· Albert Gao, Bing Xue, Andrea Zanette·· 3 小时前

SLVR:用类人推理流程实现结构化潜在视觉推理

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

AI 导读

研究者提出训练框架 SLVR,将多模态推理组织为规划、定位、证据选择与推理整合等类型化潜在阶段,避免推理时生成文本思维链。基于 Qwen2.5-VL-7B,SLVR 在 MMVP 上绝对提升 +9.4、BLINK Relation 上提升 +14.2,并在 V*、MathVista、ChartQA 上均有改进,该工作已被 NeurIPS 接收。

正文

View PDF HTML (experimental)

Abstract:Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection. We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and latent reasoning by organizing multimodal reasoning into typed latent stages for planning, grounding, evidence selection, and reasoning integration.
SLVR first trains the model to rely on the image by masking answer-revealing text and contrasting the correct answer with visually plausible distractors. It then organizes reasoning into latent stages for planning, grounding, evidence selection, and integration, supervising each stage with the corresponding signal: plans, boxes, visual evidence, and final rationales. This gives latent reasoning an explicit functional structure while avoiding generated textual chains at inference time.
Built on Qwen2.5-VL-7B, SLVR improves consistently across multimodal reasoning benchmarks, with absolute gains of +9.4 on MMVP and +14.2 on BLINK Relation, as well as improvements on V*, MathVista, and ChartQA. These results suggest that structured latent supervision can improve fine-grained visual reasoning without the decoding overhead of textual CoT. Project page is available \href{this https URL}{here}.
Comments: Accepted by NeurIPS this http URL page \href{this https URL}
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.10563 [cs.CV]
  (or arXiv:2610.10563v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.10563

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bo Gao [view email]
[v1] Thu, 1 Oct 2026 23:04:31 UTC (1,907 KB)

来源:arXiv:cs.AI · arxiv.org