arXiv:cs.CL· Wenyu Huang, Xinyu Hou, Pavlos Vougiouklis, Ruofei Lai, Jeff Z. Pan·· 3 小时前AI 评分33
超越结果奖励:为搜索智能体构建与分配检索信用
Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
AI 导读
针对 RLVR 依赖稀疏结果监督、信用分配困难的问题,研究者系统考察了从中间检索步骤提供学习信号的多种奖励塑形与信用分配策略。基于此提出一个将中间信号与最终结果奖励结合的训练框架,在多个 benchmark 上、匹配训练条件下提升了搜索智能体整体表现,并显示中间信号的选择及其信用分配位置都会影响训练行为。
正文
Abstract:Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10179 [cs.CL] |
| (or arXiv:2610.10179v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10179 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wenyu Huang [view email]
[v1]
Wed, 7 Oct 2026 14:48:47 UTC (167 KB)
来源:arXiv:cs.CL · arxiv.org