跳到正文
arXiv:cs.AI· Yu Li, Zheng Zhang, Xin Liu, Shengtian Yang, Guangfeng Cai, Lei Feng·· 5 小时前AI 评分40

CITA:面向长程工具调用智能体的比较式价值估计

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

AI 导读

针对长程工具调用中最终结果奖励难以进行信用分配的问题,研究者提出 Comparative Inference for Tool-use Agents(CITA),训练 Comparative Inference Model(CIM)在执行下一次工具调用前估计其长期价值。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.
Comments: NeurIPS 2026 Poster
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02330 [cs.AI]
  (or arXiv:2610.02330v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02330

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yu Li [view email]
[v1] Thu, 1 Oct 2026 18:04:04 UTC (982 KB)

来源:arXiv:cs.AI · arxiv.org