arXiv:cs.AI· Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Herv\'e Muyal, Marcelo Yannuzzi·· 5 小时前AI 评分33
面向基于 SLM 的智能体任务-工具意图匹配
Toward SLM-based agentic task-tool intent matching
AI 导读
研究用小型语言模型(SLM)充当任务-工具相关性分类器,对智能体每次工具调用独立评估其是否契合任务意图,并输出相关性信号供下游执行。团队构建了所需工具跨不同 MCP 服务器的多工具任务数据集,通过提示词优化、监督微调和 GRPO 强化学习对 SLM 进行优化与专门化。
正文
Abstract:Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent's underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task's intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03213 [cs.AI] |
| (or arXiv:2610.03213v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03213 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chiara Troiani [view email]
[v1]
Fri, 2 Oct 2026 12:31:01 UTC (69 KB)
来源:arXiv:cs.AI · arxiv.org