重新思考潜在视觉推理:让潜在推理锚定视觉证据
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
针对潜在视觉推理(LVR)中潜在 token 对改变正确答案的图像扰动响应微弱这一"潜在证据信用缺口",研究者提出 ReaLVR,将视觉证据监督引入模型自身的潜在推理轨迹。该方法在三个模型家族上均优于现有 LVR 基线,在 Qwen2.5-VL-7B 上取得五项任务平均 63.7% 的最高成绩,并首次将潜在空间视觉推理扩展至 235B 规模。
Published on Sep 28
Authors:
,
,
,
,
,
,
,
,
,
,
,
Abstract
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
View arXiv page View PDF Project page GitHub 7 Add to collection
Get this paper in your agent:
hf papers read 2609.34563
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.34563 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.34563 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.34563 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co