arXiv:cs.CL· Sergei Polezhaev, Barys Liskavets, Ori Press, Alexander Golubev·· 3 小时前AI 评分46
从任务结果训练 LLM 智能体的顾问模型 Caddie
Training Advisors for LLM Agents from Task Outcomes
AI 导读
研究者提出 Caddie,一种通过任务最终成败结果训练 critic 为 LLM 智能体提供自然语言分析与建议的方法,训练时冻结基座模型。用单一基座模型在 multi-hop 问答上训练出的 Qwen3-4B critic,在 MuSiQue 上把 Qwen3-4B 成功率提升逾 25 个百分点,超过无 critic 的 Kimi K3,并在 τ³、DeepDive 等域外基准上零训练迁移获益。
正文
Abstract:Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $\tau^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.
| Comments: | 26 pages, 11 figures |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09858 [cs.CL] |
| (or arXiv:2610.09858v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09858 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sergei Polezhaev [view email]
[v1]
Wed, 7 Oct 2026 11:14:26 UTC (2,328 KB)
来源:arXiv:cs.CL · arxiv.org