arXiv:cs.AI· Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang·· 3 小时前
VICO:面向视觉语言模型推理的视觉环境协同演化框架
VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
AI 导读
马里兰大学等机构提出 VICO 协同演化框架,让视觉语言模型的训练环境随模型能力同步进化。其 EnvRewriter 通过编辑场景图、图表表格等图像结构并按通过率奖励校准难度,无需额外人工标注。VICO-8B 在九个多模态基准上域外任务最高提升 5.0%,比最强自演化与文本编辑协同基线分别高 4.3% 和 8.4%,且标注样本用量比图表专用 RLVR 方法少 16-160 倍。
正文
Authors:Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang
Abstract:Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many
become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the
visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor
and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as
scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty
is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with
actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and
visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest
self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized
RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution,
VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10782 [cs.CV] |
| (or arXiv:2610.10782v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10782 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Meng Lu [view email]
[v1]
Wed, 7 Oct 2026 18:39:32 UTC (22,719 KB)
来源:arXiv:cs.AI · arxiv.org