arXiv:cs.LG(机器学习,全量分类)· Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang·· 14 小时前AI 评分39
基于选择的结构化推理 SSR:面向高效多模态搜索智能体
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
AI 导读
研究者提出 Selection-based Structured Reasoning(SSR),将多模态智能体的推理从开放式生成改为从预设自然语言候选集中选择,无需辅助任务头,并借助共享上下文 KV cache 并行打分。在 2B 和 4B 模型、七个多模态搜索基准上,SSR 任务表现与同规模领先搜索智能体相当,每轮推理延迟降低超 90%,单题模型总推理延迟降低 28-54%。
正文
Authors:Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
Abstract:Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.01892 [cs.LG] |
| (or arXiv:2610.01892v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01892 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaoyu Zhu [view email]
[v1]
Thu, 1 Oct 2026 15:42:14 UTC (5,441 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org