arXiv:cs.LG· Jaewan Park, Solbee Cho, Jay-Yoon Lee·· 4 小时前AI 评分32
DAC:面向LLM智能体搜索的角色分解协同训练框架
Cross-Agent Learning Signals Enable Coordinated Role-Decomposed LLM Training
AI 导读
研究者提出 DAC(Divide and Cooperate)角色分解训练框架,在给定任务专属外部验证信号的前提下,用角色专属交叉验证奖励分别训练 searcher 和 generator。
正文
Abstract:Agentic search systems must coordinate evidence acquisition and response generation, yet existing approaches either couple both roles under a single agent objective or decompose them without disentangling their respective contributions to the final outcome. We introduce DAC (Divide and Cooperate), a role-decomposed training framework that, given task-specific external verification signals, trains a searcher and a generator with role-specific cross-verification rewards. DAC allows the generator to abstain when the retrieved evidence appears insufficient, and uses this decision together with externally evaluated search sufficiency to assign appropriate credit to each role. To prevent degenerate over-abstention, we further introduce hard-positive evidence augmentation, which discourages abstaining on sufficient but hard evidence. Across seven general and multi-hop QA benchmarks and two model backbones, DAC consistently outperforms strong single-agent and multi-agent baselines. Controlled evaluations show that these gains arise from coordinated improvements in both search and generation. Our results highlight the importance of explicitly and properly assigning credit across interacting roles when training agentic search systems.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.10684 [cs.LG] |
| (or arXiv:2606.10684v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2606.10684 arXiv-issued DOI via DataCite |
Submission history
From: Jaewan Park [view email]
[v1]
Tue, 9 Jun 2026 10:40:55 UTC (589 KB)
[v2]
Wed, 7 Oct 2026 04:10:39 UTC (1,565 KB)
来源:arXiv:cs.LG · arxiv.org