跳到正文
arXiv:cs.AI· Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang·· 3 小时前

VICO:面向视觉语言模型推理的视觉环境协同演化框架

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

AI 导读

马里兰大学等机构提出 VICO 协同演化框架,让视觉语言模型的训练环境随模型能力同步进化。其 EnvRewriter 通过编辑场景图、图表表格等图像结构并按通过率奖励校准难度,无需额外人工标注。VICO-8B 在九个多模态基准上域外任务最高提升 5.0%,比最强自演化与文本编辑协同基线分别高 4.3% 和 8.4%,且标注样本用量比图表专用 RLVR 方法少 16-160 倍。

正文

Authors:Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang

View PDF HTML (experimental)

Abstract:Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many
become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the
visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor
and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as
scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty
is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with
actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and
visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest
self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized
RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution,
VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.10782 [cs.CV]
  (or arXiv:2610.10782v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.10782

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Meng Lu [view email]
[v1] Wed, 7 Oct 2026 18:39:32 UTC (22,719 KB)

来源:arXiv:cs.AI · arxiv.org