arXiv:cs.AI· Chaoliang Yan, Zihao Xu, Yuekang Li, Shangzhi Xu, Yi Liu, Gelei Deng, Siqi Ma·· 3 小时前
编码智能体技能冲突研究:过多安装技能如何相互干扰
One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
AI 导读
研究发现近四分之一的已安装编码智能体技能会与功能相似的技能共存,相似技能在没有降低任务完成率的情况下抢走了五分之一运行次数,并导致被替代技能丧失超过三分之一的独有核心功能。冲突在首次技能读取时即已决定,安装位置决定哪个技能运行,而最终回复仅在0.9%的替代运行中提及实际使用的技能。在首次读取处设置工具前钩子可将核心功能保真度恢复到优先打开已安装技能的水平。
正文
Abstract:Coding agents are extended with agent skills, directories whose this http URL tells the model when and how to perform a task. Because skills come from independent sources (teams, developers, plugins, copied collections), an installed skill can be co-installed with a similar skill doing the same job, and the model picks between them by name and description alone. In a conflict, the installed skill loses core functions (e.g., a ban on touching git) because the similar skill runs instead or changes what it does. The task still passes, so benchmarks that check only task completion miss such cases. We present the first empirical study of such conflicts. From snapshots of 20,947 repositories, we mine 822,109 candidate similar-skill pairs, have an LLM judge a stratified sample of 3,754, and run 312 confirmed pairs on three models (6,368 runs, 169,294 tool calls, 542 agent-hours). We report five findings. (1) Conflict-prone skills are common: nearly one in four installed skills is co-installed with one that does the same job, and 37% of judged skills sit inside copied collections. (2) Most such pairs involve normative skills, then capability skills. (3) Without lowering task completion, a similar skill takes one in five runs from the installed skill, and runs that open the similar skill first lose over a third of the exclusive core functions that only the installed skill fulfills. (4) Install location decides which skill runs, listing order barely matters, and the final reply names the skill used in only 0.9% of substituted runs. (5) Conflicts are decided at the first skill read, almost always before any file is changed, and a pre-tool hook at that read restores fidelity on exclusive core functions to the level of runs that open the installed skill first. Benchmarks should thus score exclusive core functions, and platforms should guard the first read and show which skill ran.
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11647 [cs.SE] |
| (or arXiv:2610.11647v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11647 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zihao Xu [view email]
[v1]
Thu, 8 Oct 2026 10:20:47 UTC (821 KB)
来源:arXiv:cs.AI · arxiv.org