跳到正文
arXiv:cs.AI· Xin Yu, Lizhu Zhang, Jiamu Bai, Yanhong Wu, Zellux Wang, Serena Li, Weiwei Li, Lingzhou Xue, Xiangjun Fan, Bo Peng·· 6 小时前AI 评分41

MLToolBench:让智能体学会在机器学习开发中使用工具

MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

AI 导读

研究团队提出 MLToolBench,一套覆盖数据检查、代码验证与实验诊断的可执行工具及配套 SFT 与 RL 训练流程,并用 SPICE 将工具调用转化为轮级奖励。在 80 个合成任务上训练后,Qwen3-8B 的域内成功率从 24.8% 升至 52.4%,Qwen3.5-35B-A3B 从 35.6% 升至 69.2%,后者域外成绩也从 31% 提升到 48%。

正文

View PDF HTML (experimental)

Abstract:Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagnostic calls acquire evidence whose value depends on subsequent decisions, so final outcomes provide limited guidance on which calls to reinforce. We address this challenge with SPICE, which measures how privileged context changes the likelihood of a sampled tool action and uses this difference as a turn-level reward alongside the final outcome. We train on 80 synthetic tasks and evaluate on 25 in-domain and 10 out-of-domain tasks. Providing tool interfaces and descriptions alone yields inconsistent gains across unadapted models. With the same diagnostic interface, our training pipeline raises in-domain success from 24.8% to 52.4% for Qwen3-8B and from 35.6% to 69.2% for Qwen3.5-35B-A3B. The latter also improves from 31% to 48% out-of-domain, supporting learned diagnostic tool use on held-out sources and targets.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.36679 [cs.AI]
  (or arXiv:2609.36679v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.36679

arXiv-issued DOI via DataCite

Submission history

From: Xin Yu [view email]
[v1] Tue, 29 Sep 2026 04:17:26 UTC (3,542 KB)
[v2] Tue, 6 Oct 2026 06:57:34 UTC (3,543 KB)

来源:arXiv:cs.AI · arxiv.org