arXiv:cs.AI· Suxin Ji, Hungtao Wan, Mingjun Liu, An Zhang·· 6 小时前AI 评分35
Selective-RL:将 RL 训练阶段更新选择性迁移至视觉推理模型
Selective Transfer of RL Updates for Visual Reasoning
AI 导读
研究者提出 Selective-RL,通过隔离强化学习阶段产生的参数更新、保留其主导矩阵方向并迁移至 VLM 的语言模块,实现跨模型能力迁移。在三个模型家族、五个视觉推理基准的 15 项对比中,该方法有 12 项优于完整更新插值,其中 Qwen 接收模型在 MathVision 上提升 8.55 个百分点。对照实验显示,仅靠更新幅度或任意低秩无法复现该增益,代码已开源。
正文
Abstract:Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08659 [cs.CV] |
| (or arXiv:2610.08659v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08659 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Suxin Ji [view email]
[v1]
Tue, 6 Oct 2026 16:43:55 UTC (2,711 KB)
来源:arXiv:cs.AI · arxiv.org