跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Haocun Ye, Xinlong Jiang, Qile Chen, Bingyu Wang, Teng Zhang, Shubai Chen, Tingyu Wu, Zhenkun Zheng, Yiqiang Chen·· 14 小时前AI 评分49

相同奖励,不同技能:多模态 RL 何时学会“看”

Same Reward, Different Skills: When Multimodal RL Learns to Look

AI 导读

多模态 RLVR 即使训练时不提供视觉信息也能提升视觉语言基准分数,3B 模型在测试时用图像可恢复约一半的真实图像增益,7B 模型恢复近五分之四。研究提出“视觉可解性”设计规则,在反事实坐标场景中用标准 GRPO 训练 7B 模型,将目标发现准确率从 0.425 提升至 0.875,且泛化到未训练过的问题类型。将测试图像替换为灰色画布后发现率降为零,说明奖励要求什么,RL 就学什么。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01908 [cs.LG]
  (or arXiv:2610.01908v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01908

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haocun Ye [view email]
[v1] Thu, 1 Oct 2026 15:53:22 UTC (2,485 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org