跳到正文
arXiv:cs.CL· Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Qimin Wu, Jingru Fan, Chen Qian, Yanfeng Wang, Yu Wang·· 4 小时前AI 评分42

以文本为中心的后训练实现全模态推理

Text-Centric Post-Training for Omni-Modal Reasoning

AI 导读

研究提出以文本为中心的后训练范式:纯文本训练承担主要推理优化,再用减少数据的原生音视频 RL 精修感知。最佳纯文本配置下,SFT 加 RL 将 Qwen2.5-Omni-7B 九项推理得分的几何均值较基座模型提升 25.83%,GPU 小时数比完整原生音视频路线少 56.6%。

正文

View PDF HTML (experimental)

Abstract:Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.02819 [cs.CL]
  (or arXiv:2610.02819v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02819

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ziyang Cheng [view email]
[v1] Fri, 2 Oct 2026 05:08:25 UTC (14,383 KB)

来源:arXiv:cs.CL · arxiv.org