跳到正文
arXiv:cs.CL· Arkajyoti Chakraborty, Aryan Tayal, Ishika Agarwal, Tanner Sorensen, Justin Chiu, Alessandro Di Bari, Neha Gupta, Andreas Stolcke·· 3 小时前AI 评分38

ToolRACER:用于智能体训练与评估的鲁棒智能体对话仿真资源

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

AI 导读

ToolRACER 是一套合成数据生成流程,协同用户、助手与工具仿真模型生成并验证多轮人机交互,构建出覆盖六个领域、55 种 persona、5.6K 条对话轨迹的 ToolRACERBench,其中约 66% 含易失败场景。在 ToolRACERBench 上训练的模型在 τ²-bench 与 ACEBench 上的端到端智能体准确率均有提升,在小型语言模型中与域内数据集混合时增益显著。

正文

View PDF HTML (experimental)

Abstract:Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66\% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as $\tau^2$-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across $\tau^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09163 [cs.CL]
  (or arXiv:2610.09163v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09163

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Arkajyoti Chakraborty [view email]
[v1] Tue, 6 Oct 2026 22:06:17 UTC (1,424 KB)

来源:arXiv:cs.CL · arxiv.org